Updated · By Pete Bromfield, IB examiner

IA statistics, step by step · χ² goodness of fit

The χ² goodness-of-fit test, step by step

AI SLAI HL

Do lunchtime library visits spread evenly over the week? The goodness-of-fit test compares observed counts with the counts a model predicts. Every expected frequency, contribution and the decision are worked here.

Example data, invented for this guide. The context is realistic, but the numbers were made up to show the method. Use your own collected or sourced data in your IA.

When to use it

  • Counts in categories of one variable.
  • A clear model for the expected proportions: equal, known proportions (days in each month, national figures) or a distribution.
  • Every expected frequency at least 5.

Course fit: AI SL and HL: the χ² goodness-of-fit test is in the AI SL core. Fitting a binomial, Poisson or normal model and losing a degree of freedom per estimated parameter is covered on the binomial and Poisson page. Not in the AA courses.

The example data

The school library counted lunchtime visitors (one visit per student per day) for one ordinary week: 200 visits in total.

Example data: observed counts
DayCount
Monday38
Tuesday45
Wednesday52
Thursday41
Friday24

Equal proportions: is every weekday the same?

If visits were spread evenly, each day would expect one fifth of the 200 visits.

Step 1 · Hypotheses and expected frequencies

H0: the categories are equally likely. H1: they do not. Significance level 5%.

E = N × probability,   N = 200

Step 2 · Compare observed and expected

Observed and expected frequencies
DayObserved OProbabilityExpected E(O − E)² / E
Monday380.200040.000.1000
Tuesday450.200040.000.6250
Wednesday520.200040.003.600
Thursday410.200040.000.02500
Friday240.200040.006.400
Σ20010.75

Smallest expected frequency: 40.00 (0 below 5, so the test can be used).

Observed and expected frequencies for day
Observed counts beside the counts the model expects. Where are the biggest gaps?

In the full worked analysis

  • The rest of the working: steps 3 to 3
  • What the example shows, in context
  • On a GDC: TI-84 Plus CE, TI-Nspire CX and Casio fx-CG50
  • What examiners look for
  • Common mistakes
  • Limitations to discuss

How this maps to the IA criteria

These are the current criteria A–E, for exams up to November 2028. For the new courses (first assessment May 2029) the IB has confirmed one set of four criteria for SL and HL: A Problem specification (4 marks), B Abstraction (6), C Computation (4) and D Interpretation (6), still 20 marks and 20% of the grade at both levels — see the IB's new AA and AI subject briefs. The detailed descriptors come with the new guide; check with your teacher which criteria apply to you. The advice is our summary, not the IB's wording.

Get feedback on your write-up Self-check your draft

Frequently asked questions

What is the difference between goodness of fit and the test for independence?

Goodness of fit compares one variable's counts with a model's predictions. The test for independence asks whether two categorical variables in a two-way table are linked.

How many degrees of freedom does a goodness-of-fit test have?

ν = (number of categories) − 1 − (number of parameters estimated from the data), counted after any combining.

Can I use a goodness-of-fit test for a normal distribution?

Yes: group the data into classes, find expected counts from the normal model, combine classes so each expected count is at least 5, and subtract two degrees of freedom if you estimated both the mean and the standard deviation.

Next steps

Related: Binomial and Poisson models · Chi-squared test for independence. Or analyse your own data, find a data set in the IA data bank, and see what the IA package adds.

Free: the IA checklist an examiner uses

Every check for Criteria A–E in a 4-page PDF, the mistakes that cost the most marks and a self-assessment grid. We'll email it with a short IA tip every few days, timed to your deadline if you give it. Free — no account, no payment.