IA statistics, step by step · Sampling and collecting data
Sampling and collecting data for a statistics IA: bias, size and ethics
The best analysis cannot rescue a biased sample. This page shows how to choose a sampling method, works out a stratified sample, and lists what to plan before you collect a single value — including consent if you survey people.
Example data, invented for this guide. The context is realistic, but the numbers were made up to show the method. Use your own collected or sourced data in your IA.
When to use it
- Before you collect anything.
- When you use secondary data: how was it sampled, by whom, and for what purpose?
- When you write the method section of your IA.
Course fit: Every course: populations and samples, sampling techniques (simple random, systematic, stratified, quota and convenience), reliability of data sources and bias are in the SL core of both AA and AI.
The example data
A school of 1000 students: a survey of 60 students, stratified by year group.
| Year group | Students |
|---|---|
| Year 7 | 180 |
| Year 8 | 175 |
| Year 9 | 160 |
| Year 10 | 150 |
| Year 11 | 140 |
| Year 12 | 95 |
| Year 13 | 100 |
A stratified sample, worked
To survey 60 of the 1000 students so that every year group is represented in proportion, split the sample in the same proportions as the school.
Proportional allocation
sample from a stratum = (stratum size / population size) × sample size = Ni / 1000 × 60
| Year group | Size Ni | Exact share | Sample (rounded) |
|---|---|---|---|
| Year 7 | 180 | 10.80 | 11 |
| Year 8 | 175 | 10.50 | 11 |
| Year 9 | 160 | 9.600 | 10 |
| Year 10 | 150 | 9.000 | 9 |
| Year 11 | 140 | 8.400 | 8 |
| Year 12 | 95 | 5.700 | 6 |
| Year 13 | 100 | 6.000 | 6 |
| Total | 1000 | 61 |
The rounded sizes add to 61, not 60: adjust the stratum closest to a .5 and say so.
Sampling methods compared
| Method | How it works | Example |
|---|---|---|
| Simple random | Every member of the population has an equal chance, chosen with random numbers. Needs a full list of the population. | Your class list, a register of matches, a database of products. |
| Systematic | Every k-th member from a random start. Easy, but it can line up with a pattern in the list. | Every 5th customer, every 10th line of a text. |
| Stratified | Split the population into groups (strata) and sample each in proportion. Guarantees each group is represented. | Year groups in a school, regions of a country. |
| Quota | Fill a target number in each group, but not at random within it. Quicker, open to bias. | Ask the first 10 people from each year group you meet. |
| Convenience | Whoever is easiest to reach. Fast, usually biased: say how, and what it does to your conclusion. | Your friends, your class, your own social media followers. |
Ethics and consent
- Ask your teacher what your school requires before surveying anyone; under-18s may need a parent's or the school's permission.
- Tell participants what the data are for, that taking part is voluntary and that they can withdraw.
- Collect only what you need; keep responses anonymous and store them securely.
- Never put names or identifying details in your IA; anonymise any raw data in the appendix.
- Avoid questions that could upset people (health, weight, family) unless your teacher has agreed and you have a clear reason.
What the example shows
- Year 7 is 180 of 1000 students, so it gets 10.8 of the 60 places, rounded to 11.
- The rounded sizes add to 61: one too many. Remove one from Year 8 (exact share 10.5, the one that was rounded up from exactly .5) and say so in your method.
- Within each year group, choose students by simple random sampling (number them and use a random number generator) so the sample is random as well as proportional.
This is our interpretation of invented example data, to show the kind of thinking examiners reward. Your interpretation must be your own, about your own data.
On a GDC
Enter the data first, then use the menu below. Menus differ slightly between operating systems; check your calculator's manual.
- TI-84 Plus CE: [stat] → EDIT, data in L1; [stat] → CALC → 1-Var Stats. It gives x̄, Σx, Sx, σx, n, minX, Q1, Med, Q3 and maxX.
- TI-Nspire CX: Lists & Spreadsheet, then menu → Statistics → Stat Calculations → One-Variable Statistics.
- Casio fx-CG50: STAT mode, data in List 1; CALC → 1-VAR (check SET points at List 1).
Show one calculation by hand in your IA so the examiner sees you understand it, then say which technology did the rest.
What examiners look for
- The population is defined, and the sample's size and method are justified — not just stated.
- Possible bias is named for this sample (who is missing, who is over-represented) and its effect on the conclusion discussed.
- Secondary data are referenced in full, with a comment on how reliable the source is.
- Ethics: consent, anonymity and the right to withdraw if people are involved.
Common mistakes
- Calling a convenience sample “random”.
- A sample size chosen with no reason (or too small for the test to be valid).
- Surveying your friends about something they share with you.
- No mention of how the data were checked for errors.
- Collecting personal data (names, dates of birth) you do not need.
Limitations to discuss
- A larger sample reduces random error but not bias.
- Self-reported data (screen time, sleep, revision hours) are often inaccurate in a predictable direction.
- Secondary data were collected for someone else's purpose; definitions may not match your question.
How this maps to the IA criteria
- A Criterion A (Presentation): A method section a reader can follow: population, sample, method, size and timing, in that order.
- B Criterion B (Mathematical communication): Variables defined with units and how they were measured; tables of raw data labelled and in an appendix.
- C Criterion C (Personal engagement): Decisions that are your own — why this population, this method, this size — and data you collected or found for a reason.
- D Criterion D (Reflection): Honest discussion of bias, reliability and what the sample lets you conclude about the population.
- E Criterion E (Use of mathematics): Sampling calculations (proportional allocation, sample sizes) done correctly; data fit for the techniques you plan.
These are the current criteria A–E, for exams up to November 2028. For the new courses (first assessment May 2029) the IB has confirmed one set of four criteria for SL and HL: A Problem specification (4 marks), B Abstraction (6), C Computation (4) and D Interpretation (6), still 20 marks and 20% of the grade at both levels — see the IB's new AA and AI subject briefs. The detailed descriptors come with the new guide; check with your teacher which criteria apply to you. The advice is our summary, not the IB's wording.
Frequently asked questions
How big should my sample be for a Maths IA?
Big enough for your technique to be valid and your conclusion to be worth something: every expected frequency at least 5 for χ², and usually at least 20–30 values per group for a t-test or a normal model. Explain your choice and what limited it.
Can I use secondary data in my IA?
Yes. Reference the source, explain how it was collected, check it for errors and comment on its reliability. Secondary data still need a sampling decision: which part of it you use and why.
Do I need consent for a survey?
If you collect data from people, ask for their consent, explain what it is for, keep it anonymous and let them withdraw. Check your school's rules with your teacher first.
Next steps
- Criterion E: use of mathematicsWhat “commensurate with the level of the course” means for statistics, at SL and HL.
- Criterion D: reflectionSample, bias, assumptions and causation: where statistics IAs gain or lose marks.
- Plan your statistics IAThe section-by-section framework for a statistics exploration, with your own notes.
- Get feedback on your write-upCriterion-by-criterion feedback on your draft, with evidence from your own text.
Related: Choosing the right test · Cleaning data and outliers. Or analyse your own data, find a data set in the IA data bank, and see what the IA package adds.
Free: the IA checklist an examiner uses
Every check for Criteria A–E in a 4-page PDF, the mistakes that cost the most marks and a self-assessment grid. We'll email it with a short IA tip every few days, timed to your deadline if you give it. Free — no account, no payment.
While you wait for the email: read the full statistics workflow →