IA idea · Survey-based investigations
How big should a school poll be? Sampling from a known population
Research question
By taking repeated samples of different sizes from a dataset where the true answer is known, how does the margin of error shrink with sample size, and how many students would a school poll need for a margin of ±5 percentage points?
Adapt it: change the place, the data or the comparison until the question is yours.
Free: the A–E checklist an examiner uses, by email ↓
Why it makes a good exploration
It makes sampling error visible: you see the spread of estimates around a known truth before you trust your own single poll.
The mathematics you'll need
- Random sampling methods
- Sampling distribution of a proportion
- Standard error √(p(1 − p)/n)
- Margin of error and sample size
- Comparing random, convenience and stratified samples
Course labels show where a technique sits; using maths from outside your course is fine if you explain it clearly and say it is new to you.
Where the data comes from
Use a public dataset as a 'population' (the UCI repository has a student-performance dataset), then run your own small poll.
- UCI Machine Learning Repository — 600+ classic, documented datasets (wine quality, student performance, bike sharing…).
- Desmos graphing calculator — Free graphing and regression (y₁ ~ ax₁ + b) — fit models to your data and show residuals.
Cite every source in a footnote where you use it and in your bibliography. Check the licence of any dataset you download.
A possible outline
- Choose a population dataset and a yes/no variable.
- Take many samples of each size and record the estimates.
- Compare the spread with the standard-error formula.
- Compare sampling methods.
- Plan and run your own poll and reflect.
Pitfalls that cost marks
- Too few repeated samples.
- Treating a convenience sample as random.
- Using the formula without checking it against the simulation.
Showing personal engagement
- Run a real poll on a school question.
- Predict the needed sample size first.
- Test a convenience sample against a random one.
See Criterion C: personal engagement for what examiners look for.
Which course is it for?
| Course | Fit | Maths to lean on |
|---|---|---|
| AA SL | Good fit | Random sampling methods; Sampling distribution of a proportion |
| AA HL | Not a natural fit | The mathematics is mainly from the AI course; at AA HL the exploration would need an AA-level approach (calculus, proof or probability theory) to reach the top of Criterion E. |
| AI SL | Good fit | Random sampling methods; Sampling distribution of a proportion |
| AI HL | Good fit | Random sampling methods; Sampling distribution of a proportion |
Level: Solid. Needs some independent work beyond class examples. See how the IA differs between AA and AI, SL and HL.
How this idea reaches the top bands
Personal engagement (C)
Design every part yourself: the question, the sampling frame, the wording and the analysis plan. Piloting the questionnaire and changing it is engagement an examiner can see.
Reflection (D)
Reflect on bias (who answered, who didn't, how wording steered answers) and on what a significant result can and can't show about the whole population. For this idea, start with: too few repeated samples — say how it affects your answer.
Use of mathematics (E)
SL: A sampling method justified, a sample size worked out from the test you plan, the test's conditions checked, one calculation shown and the p-value interpreted in context.
HL: Confidence intervals, a justified choice between tests, a test of a model's fit, or a statistical estimate of how much bias could change the conclusion.
Criteria A and B (presentation and communication) work the same way for every idea: see the guides to Criterion A and Criterion B.
Taking it further
Adjust for the finite population of your school and see how much it matters.
Extending it for HL
Add confidence intervals for the proportions you report, and estimate how large a non-response bias would have to be to overturn your result.
See a complete IA, marked
Our annotated exemplar Do students who sleep less react more slowly? (AI SL) asks a different question, but shows how a complete surveys exploration is structured and marked, with an examiner's comment on every criterion. Free excerpts and the full marking table are on its page.
Before you start: the checklist an examiner uses
Every check for Criteria A–E in a 4-page PDF, the mistakes that cost the most marks and a self-assessment grid. We'll email it with a short IA tip every few days, timed to your deadline if you give it. Free — no account, no payment.
While you wait for the email: read the free excerpt of a complete, annotated IA (Sleep and reactions (statistics)) →
Turn this idea into your IA
Similar ideas
- Counting a hidden population: capture–recapture at schoolAI SLAA SLAI HLAA HLSolid
- Does the wording change the answer? A split-sample survey experimentAI SLAI HLAA SLSolid
- Do we know how long we spend on our phones? Estimates versus screen-time dataAI SLAA SLAI HLAccessible
- Do students who sleep less react more slowly?AI SLAI HLAA SLAA HLAccessible
All surveys ideas · AI SL ideas · AI HL ideas · AA SL ideas · All 239 IA ideas