Top-band exemplar · AI SL · written to a full-marks standard
Is my 07:40 bus late because of the rain?
Written by IB Math Revision to show a top-band (full-marks) standard — not a real student's IA, not moderated by the IB; marks can't be guaranteed. The marks below are our examiner-style judgement against the criteria. This exploration was never submitted in an IB session and we do not claim it scored 20/20.
An AI SL statistics exploration written to a top-band (full-marks) standard, built around a real decision: a normal model fixed in advance from a pilot log and tested with χ² goodness of fit, a χ² test of independence redesigned when the first table failed the expected-frequency condition, a pooled t-test calculated by hand, a lurking variable found in a regression, and a decision made with normal probabilities.
Why it reaches the top band, criterion by criterion
| Criterion | Mark | Why |
|---|---|---|
| A · Presentation | 4/4 | Organised around a four-part aim and a decision; each section answers one part, and the conclusion answers them in order. Concise, with nothing included for its own sake. |
| B · Mathematical communication | 4/4 | Variables and the relation A = −20 + D + J defined; hypotheses, significance levels, degrees of freedom and conclusions stated in context; tables and graphs chosen to answer the questions. |
| C · Personal engagement | 3/3 | Outstanding: a real decision with a personal cost, the student's own data and categories, and choices (continuity boundary, weather recorded at 07:30, three categories) that are clearly theirs. |
| D · Reflection | 3/3 | Substantial critical reflection throughout: tests are questioned (power, validity conditions), a lurking variable is found and tested, model weaknesses are quantified and used in the decision. |
| E · Use of mathematics | 6/6 | Relevant AI SL statistics used with sophistication and rigour: each test chosen for a stated question, validity conditions checked and acted on (classes merged, a 2 × 2 table rejected, the normal model's parameters fixed in advance from a separate pilot log so that the goodness-of-fit test is valid as the SL course states it), key statistics calculated by hand and matched to the GDC, and every result interpreted in context and used in the decision. |
| Total | 20/20 | Full marks are justified because rigour and reflection run through every section: the statistics are not only correct but chosen, checked and questioned, and they answer a question that matters to the student. At SL, sophistication means this kind of judgement, not mathematics from outside the course — every technique used is in the AI SL syllabus. |
Marks are our judgement of this teaching exemplar against the current criteria, explained criterion by criterion. They are not IB moderation results, and a real IA written to this standard could still be marked differently by a teacher or moderator.
What would lose marks here
The same exploration, with these changes, would drop out of the top band:
- E Running the χ² test on the 2 × 2 table anyway, with an expected frequency below 5, would be an error in the use of the test and would cap E.
- E Taking the normal model's mean and standard deviation from the same 64 days and then running the goodness-of-fit test as if the model had been given in advance — the SL test needs a distribution fixed before the data are seen, which is why the pilot log's figures are used.
- E Pasting GDC output (“p = 0.0000…”) without hypotheses, degrees of freedom or a conclusion in context would count as a routine use of technology.
- D Concluding “delays cause slower journeys” from r = 0.35 without checking for the weather as a lurking variable.
- C A topic chosen from a list with data downloaded but no decisions of the student's own — the same tests on someone else's data would earn much less for C.
- A Adding tests that do not serve the aim (for example a regression of arrival time on day of the week “for completeness”).
Excerpts with examiner annotations
Free sections are shown below with comments; the rest is in the full exemplar, available in the protected viewer with the IA package or a Pro plan.
1. Introduction
My school gives a late mark to anyone who arrives after registration at 08:35, and three late marks in a term mean an after-school detention. I take the 07:40 bus, which is timetabled to reach the stop outside school at 08:18, and from there it is a ten-minute walk to my form room, so I need to be off the bus by 08:25. Last year I got four late marks, all — as far as I remember — on rainy days. My mother's view is that “the bus is only ever late when it rains” and that I should take the earlier 07:25 bus on wet mornings; I would rather have fifteen more minutes in bed.
Aim. Using a log of 64 journeys, I will (1) describe how reliable the 07:40 bus is and decide whether a normal distribution is a reasonable model for its arrival time, (2) test whether wet weather and late arrival are associated, (3) find out whether rain makes the journey itself slower or only delays the bus leaving, and (4) use the models to decide which bus I should take on wet and dry days.
The decision in part (4) is the reason for the whole exploration, so I will judge every result by whether it helps me make it.
C A genuine, low-stakes but real decision with a personal cost on each side. The student has a stake in the answer — which is what makes the later choices look like their own.
A A numbered aim whose parts are answered in order, and a stated purpose that keeps the exploration focused.
2. Collecting the data
From the second week of September to the middle of December I recorded every school-day journey on the 07:40 bus: 64 days in all (I was ill on two days and one day the bus did not run). For each journey I noted:
\(D\) = departure delay: minutes between 07:40 and the time the bus actually left my stop, from the operator's live-tracking app (to the nearest minute);
\(J\) = journey time: minutes from leaving my stop to arriving at the school stop, from the same app;
\(A\) = arrival time at the school stop, in minutes after 08:00, so \(A = -20 + D + J\);
weather: wet if it was raining at my stop at 07:30, otherwise dry.
I chose to record the weather myself at 07:30 rather than use a daily rainfall total, because a day with rain in the afternoon is a dry morning for my bus. I defined late as \(A \ge 26\) (after 08:25), which in a continuous model means \(A > 25.5\), because the app rounds to the nearest minute. Taking 25.5 rather than 25 matters: with a standard deviation of about 3 minutes, half a minute changes a probability of being late by several percentage points.
Before this, over the last six weeks of the summer term, I had kept a pilot log of 30 journeys on the same bus, mainly to learn how the app works (Appendix D). I did not record the weather in it, so I use it only once: in Section 4, as a normal model fixed before the main data were collected.
Data note: the journey log in this exemplar was simulated by IB Math Revision to behave like a real one. A student must collect and use their own data.
| Day | Weather at 07:30 | Departed | Departure delay D (min) | Journey time J (min) | Arrived |
|---|---|---|---|---|---|
| 1 | wet | 07:45 | 5 | 45 | 08:30 |
| 2 | dry | 07:41 | 1 | 34 | 08:15 |
| 3 | dry | 07:41 | 1 | 36 | 08:17 |
| 4 | wet | 07:43 | 3 | 40 | 08:23 |
| 5 | dry | 07:46 | 6 | 38 | 08:24 |
| 6 | dry | 07:42 | 2 | 40 | 08:22 |
| 7 | dry | 07:41 | 1 | 41 | 08:22 |
| 8 | dry | 07:40 | 0 | 42 | 08:22 |
Two limitations were clear before I started. My observations are 64 consecutive school days in one autumn, so they may not represent winter snow or summer traffic, and the days are not truly independent: one week had roadworks on the main road, which could make several days in a row slow together. I kept a note of the roadworks days (Appendix C) so that I could look at them separately if they stood out.
B Variables are defined with units, and the relationship A = −20 + D + J is written down so the reader can check any row of the table.
E The continuity boundary (25.5, not 25) is a small detail that shows real understanding of what a continuous model of rounded data means.
D Limitations of the data are identified before the analysis and then checked — reflection that happens during the work, not after it.
3. How reliable is the bus?
In the full exemplar (about 2 pages). Summary statistics and box plots for dry and wet days, and a decision about an outlier. Open in the protected viewer
4. Is a normal model reasonable?
In the full exemplar (about 2 pages). A χ² goodness-of-fit test of a normal model fixed in advance from a pilot log, calculated by hand: merging classes, degrees of freedom, and what the result does and does not show. Open in the protected viewer
5. Is wet weather associated with arriving late?
A \(\chi^2\) test of independence. My first plan was a 2 × 2 table of wet/dry against late/not late. But only 10 days were late, and the expected frequency for “wet and late” would have been \(\frac{24 \times 10}{64} = 3.75\), below 5, so the test would not have been valid. Instead I used three arrival categories whose boundaries mean something to me: by 08:18 (on time), 08:19–08:23 (late bus, but I still make registration comfortably), and 08:24 or later (I have to run, or I am late).
\(H_0\): arrival category is independent of the weather. \(H_1\): arrival category is not independent of the weather. Significance level 5%.
| by 08:18 (on time or early) | 08:19–08:23 | 08:24 or later | Total | |
|---|---|---|---|---|
| Wet: observed | 0 | 11 | 13 | 24 |
| Wet: expected | 6.38 | 11.25 | 6.38 | |
| Dry: observed | 17 | 19 | 4 | 40 |
| Dry: expected | 10.62 | 18.75 | 10.62 | |
| Total | 17 | 30 | 17 | 64 |
Each expected frequency is \(\frac{\text{row total} \times \text{column total}}{64}\); for example the expected number of wet days with arrival by 08:18 is \(\frac{24 \times 17}{64} = 6.375\), and its contribution to the statistic is \(\frac{(0 - 6.375)^2}{6.375} = 6.375\). All six expected frequencies are at least 5. Adding all six contributions gives \(\chi^2_{\text{calc}} = 21.22\) with \((2-1)(3-1) = 2\) degrees of freedom, so \(p \approx 2.5 \times 10^{-5}\) (critical value 5.991).
Since \(p < 0.05\) I reject \(H_0\): there is very strong evidence that arrival time and weather are associated. The largest contributions come from the “08:24 or later” column (6.88 for wet days) and from the fact that the bus never arrived on time on a wet day in my log (0 observed against 6.38 expected).
Does rain slow the journey itself? The \(\chi^2\) test says nothing about why. To see whether the journey is slower, and not only the departure, I compared mean journey times with a one-tailed two-sample \(t\)-test. \(H_0\): \(\mu_{\text{wet}} = \mu_{\text{dry}}\); \(H_1\): \(\mu_{\text{wet}} > \mu_{\text{dry}}\). The \(t\)-test assumes that both samples come from normal distributions; the journey times in each group are roughly symmetric with no outliers, and the standard deviations (2.25 and 1.99 minutes) are similar, so I used the pooled version:
\[ s_p = \sqrt{\frac{(n_w - 1)s_w^2 + (n_d - 1)s_d^2}{n_w + n_d - 2}} = \sqrt{\frac{23 \times 2.245^2 + 39 \times 1.994^2}{62}} \approx 2.091, \]
\[ t = \frac{\bar x_w - \bar x_d}{s_p\sqrt{\frac{1}{n_w} + \frac{1}{n_d}}} = \frac{41.542 - 37.850}{2.091 \times \sqrt{\frac{1}{24} + \frac{1}{40}}} \approx 6.84, \]
with 62 degrees of freedom, giving \(p \approx 2 \times 10^{-9}\) (my GDC gave the same value). This is far below 0.05, so there is strong evidence that the mean journey time is longer on wet days — by about 3.7 minutes in my sample. So rain does not only make the bus leave late: it makes the whole journey slower, which is what I would expect with heavier traffic and passengers who take longer to board with umbrellas.
E Two tests, each chosen for a different question, with conditions checked and acted on (the 2 × 2 table is rejected because of an expected frequency below 5). Showing one expected frequency, one contribution and the pooled t calculation by hand shows understanding, not just GDC output.
D Categories are chosen for their meaning to the question, and the student is explicit that the χ² test shows association, not cause.
B Hypotheses, significance level, degrees of freedom and conclusion are stated in context for each test, using correct notation.
6. Do late buses also travel more slowly?
In the full exemplar (about 2 pages). Correlation and regression of journey time on departure delay, and how much of the relationship turns out to be the weather. Open in the protected viewer
7. Which bus should I take?
In the full exemplar (about 2.5 pages). Normal models for wet and dry days, checked against the observed proportions, a 95% threshold, and the decision with its expected number of late days per term. Open in the protected viewer
8. Conclusion and reflection
The 07:40 bus is reliable on dry mornings (median arrival 08:19, one minute after the timetable) and unreliable on wet ones (median 08:24). A \(\chi^2\) test of independence showed very strong evidence that arrival time is associated with the weather (\(p \approx 2.5 \times 10^{-5}\)), and a \(t\)-test showed that the journey itself is slower on wet days by about 3.7 minutes, not only the departure. The apparent link between late departures and slow journeys is partly explained by the weather: within wet days and within dry days it is weaker. Normal models give a probability of being late of about 0.31 on a wet day and at least 0.02 on a dry day, so my plan is to take the 07:25 bus on wet mornings, which should reduce my late marks from about 8 to about 1 a term. My mother was mostly right, though not entirely: the bus is late on dry days occasionally, too.
The main limitations are in the data rather than the methods. My log covers one autumn, so it says nothing about winter ice or the summer term; “wet” is a yes/no judgement at one moment, which cannot distinguish drizzle from a downpour; and consecutive days are not independent, so the tests may be slightly over-confident. A rainfall measurement in millimetres from a weather station near the route would let me model the effect of rain as a continuous variable, and a second term's data would show whether the conclusions hold in winter.
Doing this changed how I see the timetable: it is not a promise but the centre of a distribution, and the question worth asking is not “is the bus late?” but “how likely is it to be late enough to matter to me?”
A Every part of the aim is answered, in order, with the key figures; the personal question from the introduction is answered honestly (“mostly right, though not entirely”).
D Limitations are specific and linked to their effect (dependence makes tests over-confident) and to a concrete improvement.
Bibliography and appendices
In the full exemplar (about 1 page). Sources, technology and appendices. Open in the protected viewer
What a moderator could still ask for
- Even at this level, a moderator could ask for the uncertainty in the key probabilities — the wet-day model rests on only 24 days.
- A model that combines the weather with the departure delay, or a treatment of the dependence between consecutive days, would be a natural next step (beyond what full marks require at SL).
Frequently asked questions
Did this IA actually get full marks?
No — it is not a real student's IA and was never submitted or moderated. It was written by IB Math Revision to show a top-band (full-marks) standard; the marks are our examiner-style judgement against the criteria, and no one can guarantee a mark.
Is the data real?
No. The journey log is simulated to behave like a real one, and the exemplar says so. In your own IA the data must be yours, or a properly cited real dataset.
Is statistics enough for top marks at AI SL?
Yes, if the tests are chosen for a reason, their conditions are checked, one calculation of each kind is shown and every result is interpreted in context. What separates the top band is the reflection and the decisions, not the number of tests.
Other top-band examples
- How far can a stack of books lean over the edge of a table? (AA HL)
- What is the quickest way to deliver a leaflet to every house on my estate? (AI HL)
All 10 annotated exemplars, including a deliberately mid-band draft.
Read the whole exploration
The full “Bus lateness and rain (AI SL)” exemplar, with an examiner's note on every section, is in the IA package with the other 9 exemplars (Pro and Platinum plans include them too). €39 once, 12 months' access, 14-day money-back guarantee. It helps you write your own IA; it never writes it for you.
Already have access? Open it in the viewer · All exemplars · Want your own draft marked like this? Examiner review