IA idea · Survey-based investigations
Counting a hidden population: capture–recapture at school
Research question
Can the capture–recapture method estimate how many students use the library (or a club, or a bus) from two independent anonymous counts, and how accurate is it compared with the true number?
Adapt it: change the place, the data or the comparison until the question is yours.
Free: the A–E checklist an examiner uses, by email ↓
Why it makes a good exploration
Ecologists use this to count animals; applying it to people with a true number to compare against tests the method's assumptions directly.
The mathematics you'll need
- The Lincoln–Petersen estimator, derived from proportions
- Assumptions: independence and equal catchability
- Simulation of the estimator's spread
- Bias for small samples and a corrected estimator
- Comparing with the true count
Course labels show where a technique sits; using maths from outside your course is fine if you explain it clearly and say it is new to you.
The statistics, step by step
Worked with every number shown, with what examiners look for and the common mistakes: Chi-squared test for independence · Sampling and collecting data. Then run the same steps on your own data in Analyse my data, or start from the statistics workflow.
Where the data comes from
Two independent anonymous 'captures' (for example, tally marks with a non-identifying code students choose) on different days, plus the true count from records if available. School approval needed.
- Desmos graphing calculator — Free graphing and regression (y₁ ~ ax₁ + b) — fit models to your data and show residuals.
Cite every source in a footnote where you use it and in your bibliography. Check the licence of any dataset you download.
A possible outline
- Explain and derive the estimator.
- Design an anonymous, ethical marking method.
- Collect two samples.
- Estimate and compare with the true number.
- Simulate the estimator and reflect on assumptions.
Pitfalls that cost marks
- Identifiable codes.
- Captures that aren't independent (the same day and time).
- No check against a true value.
Showing personal engagement
- Choose a population you are part of.
- Predict the estimate's error.
- Test it first with beans in a jar.
See Criterion C: personal engagement for what examiners look for.
Which course is it for?
| Course | Fit | Maths to lean on |
|---|---|---|
| AA SL | Good fit | The Lincoln–Petersen estimator, derived from proportions; Assumptions: independence and equal catchability |
| AA HL | Good fit | The Lincoln–Petersen estimator, derived from proportions; Assumptions: independence and equal catchability |
| AI SL | Good fit | The Lincoln–Petersen estimator, derived from proportions; Assumptions: independence and equal catchability |
| AI HL | Good fit | The Lincoln–Petersen estimator, derived from proportions; Assumptions: independence and equal catchability |
Level: Solid. Needs some independent work beyond class examples. See how the IA differs between AA and AI, SL and HL.
How this idea reaches the top bands
Personal engagement (C)
Design every part yourself: the question, the sampling frame, the wording and the analysis plan. Piloting the questionnaire and changing it is engagement an examiner can see.
Reflection (D)
Reflect on bias (who answered, who didn't, how wording steered answers) and on what a significant result can and can't show about the whole population. For this idea, start with: identifiable codes — say how it affects your answer.
Use of mathematics (E)
SL: A sampling method justified, a sample size worked out from the test you plan, the test's conditions checked, one calculation shown and the p-value interpreted in context.
HL: Confidence intervals, a justified choice between tests, a test of a model's fit, or a statistical estimate of how much bias could change the conclusion.
Criteria A and B (presentation and communication) work the same way for every idea: see the guides to Criterion A and Criterion B.
Taking it further
Use three captures and compare estimators.
Extending it for HL
Add confidence intervals for the proportions you report, and estimate how large a non-response bias would have to be to overturn your result.
See a complete IA, marked
Our annotated exemplar Do students who sleep less react more slowly? (AI SL) asks a different question, but shows how a complete surveys exploration is structured and marked, with an examiner's comment on every criterion. Free excerpts and the full marking table are on its page.
Before you start: the checklist an examiner uses
Every check for Criteria A–E in a 4-page PDF, the mistakes that cost the most marks and a self-assessment grid. We'll email it with a short IA tip every few days, timed to your deadline if you give it. Free — no account, no payment.
While you wait for the email: read the free excerpt of a complete, annotated IA (Sleep and reactions (statistics)) →
Turn this idea into your IA
Similar ideas
- How big should a school poll be? Sampling from a known populationAI SLAI HLAA SLSolid
- Does the wording change the answer? A split-sample survey experimentAI SLAI HLAA SLSolid
- Getting honest answers: the randomised response techniqueAA HLAI HLAA SLAmbitious
- Estimating π by dropping needles: how fast does the estimate improve?AA HLAA SLAI SLAI HLSolid
All surveys ideas · AI SL ideas · AA SL ideas · AI HL ideas · AA HL ideas · All 239 IA ideas