IA modelling, step by step · Exponential
Exponential model, step by step: y = a e^(kx) + c and y = a·bˣ
Growth or decay by the same factor in each step — cooling, depreciation, radioactive or drug decay — is exponential. Take logs to turn it into a straight line; if the context has a non-zero asymptote (room temperature), subtract it first.
Example data, invented for this guide. The context is realistic, but the numbers were made up to show the method. Use your own measured or sourced data in your IA.
When to use it
The shape of the data
- Equal steps in x multiply y (or y − c) by roughly the same factor.
- The curve approaches a horizontal line (the asymptote) in one direction.
- ln y (or ln(y − c)) against x is close to straight.
The context
- Cooling or warming towards a surrounding temperature (Newton's law of cooling).
- Growth with a fixed percentage rate: money, early-stage populations or views.
- Decay: drug concentration, radioactive counts, the value of a car.
Course fit: Every course: exponential functions and logarithms are in AA and AI at SL; linearising with logs is AI HL; the asymptote c and least-squares refinement suit HL well.
The example data
A mug of tea cools on a desk in a room kept at 21 °C; a probe records the temperature every 5 minutes.
| i | x: Time (min) | y: Temperature of tea (°C) |
|---|---|---|
| 1 | 0 | 85 |
| 2 | 5 | 72.4 |
| 3 | 10 | 61.5 |
| 4 | 15 | 53.9 |
| 5 | 20 | 46.7 |
| 6 | 25 | 42 |
| 7 | 30 | 37.4 |
| 8 | 35 | 34.5 |
| 9 | 40 | 31.4 |
First try: y = a·e^(kx) (asymptote at 0)
Calculator exponential regression assumes the asymptote is y = 0. Fit it first — then ask whether that makes sense for tea in a 21 °C room.
Step 1 · Linearise with logarithms
Assume the horizontal asymptote is y = 0. If y = aekx + c, then ln y = kx + ln|a|: a straight line in (x, ln y) with gradient k and intercept ln|a|.
Step 2 · Tabulate the sums
| i | x | ln y | x² | x(ln y) | (ln y)² |
|---|---|---|---|---|---|
| 1 | 0 | 4.4427 | 0 | 0.0000 | 19.737 |
| 2 | 5 | 4.2822 | 25 | 21.411 | 18.337 |
| 3 | 10 | 4.1190 | 100 | 41.190 | 16.966 |
| 4 | 15 | 3.9871 | 225 | 59.807 | 15.897 |
| 5 | 20 | 3.8437 | 400 | 76.875 | 14.774 |
| 6 | 25 | 3.7377 | 625 | 93.442 | 13.970 |
| 7 | 30 | 3.6217 | 900 | 108.65 | 13.116 |
| 8 | 35 | 3.5410 | 1225 | 123.93 | 12.538 |
| 9 | 40 | 3.4468 | 1600 | 137.87 | 11.880 |
| Σ | 180.000 | 35.0219 | 5100.00 | 663.181 | 137.218 |
n = 9 points. The transformed values are rounded here; keep them unrounded in your calculator or spreadsheet.
Step 3 · Gradient
m = (nΣXY − ΣX ΣY) / (nΣX² − (ΣX)²)
m = (9 × 663.181 − 180.000 × 35.0219) / (9 × 5100.00 − 180.000²) = −335.309 / 13500 = −0.02484
Here X = x and Y = ln y.
Step 4 · Intercept
X̄ = ΣX/n = 20.000, Ȳ = ΣY/n = 3.8913
c = Ȳ − m X̄ = 3.8913 − (−0.02484) × 20.000 = 4.388
The line of best fit always passes through the mean point (X̄, Ȳ).
Step 5 · Correlation
r = (nΣXY − ΣX ΣY) / √[(nΣX² − (ΣX)²)(nΣY² − (ΣY)²)] = −335.309 / √(13500 × 8.43047) = −0.9939
r² = 0.9879: a very strong negative linear correlation between x and ln y.
Step 6 · Back-substitute
k = gradient = −0.02484; ln|a| = intercept = 4.388, so a = e4.388 = 80.49.
y = 80.49e−0.02484x
The same model with a base: y = a·bx with b = ek = 0.9755 (each unit of x multiplies y by 0.9755). Half-life: ln 2/|k| = 27.91 units of x.
This is exactly what a GDC's exponential regression does (a straight line fitted to ln y).
Step 7 · Residuals and goodness of fit (in the original units)
For each point, the residual is the observed value minus the model's value: e = y − ŷ.
| i | x | y (data) | ŷ (model) | y − ŷ | (y − ŷ)² |
|---|---|---|---|---|---|
| 1 | 0 | 85 | 80.49 | 4.51 | 20.4 |
| 2 | 5 | 72.4 | 71.09 | 1.31 | 1.73 |
| 3 | 10 | 61.5 | 62.78 | −1.28 | 1.65 |
| 4 | 15 | 53.9 | 55.45 | −1.55 | 2.41 |
| 5 | 20 | 46.7 | 48.98 | −2.28 | 5.18 |
| 6 | 25 | 42 | 43.26 | −1.26 | 1.58 |
| 7 | 30 | 37.4 | 38.20 | −0.804 | 0.647 |
| 8 | 35 | 34.5 | 33.74 | 0.758 | 0.574 |
| 9 | 40 | 31.4 | 29.80 | 1.60 | 2.55 |
| Σ | 36.70 |
- Sum of squared residuals: SSR = Σ(y − ŷ)² = 36.70
- Total sum of squares about the mean ȳ = 51.64: SST = Σ(y − ȳ)² = 2670
- Coefficient of determination: R² = 1 − SSR/SST = 1 − 36.70/2670 = 0.9863
- Root-mean-square error: RMSE = √(SSR/n) = 2.02 — a typical size of a residual, in the units of y.
r above measures how straight the transformed data are; R² here measures the fit to the original y values. They are different numbers — quote the one you mean.
Better: y = a·e^(kx) + c with c = 21 from the context
The tea cools towards room temperature, 21 °C, not 0 °C. Subtract it, linearise ln(y − 21) against x, then refine by least squares — letting c vary too, as a test of the assumption.
Step 1 · Linearise with logarithms
The horizontal asymptote is taken from the context: c = 21. If y = aekx + c, then ln(y − 21) = kx + ln|a|: a straight line in (x, ln(y − 21)) with gradient k and intercept ln|a|.
Step 2 · Tabulate the sums
| i | x | ln(y − 21) | x² | x(ln(y − 21)) | (ln(y − 21))² |
|---|---|---|---|---|---|
| 1 | 0 | 4.1589 | 0 | 0.0000 | 17.296 |
| 2 | 5 | 3.9396 | 25 | 19.698 | 15.521 |
| 3 | 10 | 3.7013 | 100 | 37.013 | 13.700 |
| 4 | 15 | 3.4935 | 225 | 52.402 | 12.204 |
| 5 | 20 | 3.2465 | 400 | 64.930 | 10.540 |
| 6 | 25 | 3.0445 | 625 | 76.113 | 9.2691 |
| 7 | 30 | 2.7973 | 900 | 83.918 | 7.8248 |
| 8 | 35 | 2.6027 | 1225 | 91.094 | 6.7740 |
| 9 | 40 | 2.3418 | 1600 | 93.672 | 5.4841 |
| Σ | 180.000 | 29.3261 | 5100.00 | 518.841 | 98.6127 |
n = 9 points. The transformed values are rounded here; keep them unrounded in your calculator or spreadsheet.
Step 3 · Gradient
m = (nΣXY − ΣX ΣY) / (nΣX² − (ΣX)²)
m = (9 × 518.841 − 180.000 × 29.3261) / (9 × 5100.00 − 180.000²) = −609.127 / 13500 = −0.04512
Here X = x and Y = ln(y − 21).
Step 4 · Intercept
X̄ = ΣX/n = 20.000, Ȳ = ΣY/n = 3.2585
c = Ȳ − m X̄ = 3.2585 − (−0.04512) × 20.000 = 4.161
The line of best fit always passes through the mean point (X̄, Ȳ).
Step 5 · Correlation
r = (nΣXY − ΣX ΣY) / √[(nΣX² − (ΣX)²)(nΣY² − (ΣY)²)] = −609.127 / √(13500 × 27.4949) = −0.9998
r² = 0.9996: a very strong negative linear correlation between x and ln(y − 21).
Step 6 · Back-substitute
k = gradient = −0.04512; ln|a| = intercept = 4.161, so a = e4.161 = 64.13.
y = 64.13e−0.04512x + 21
The same model with a base: y = a·bx + c with b = ek = 0.9559 (each unit of x multiplies y − c by 0.9559). Half-life: ln 2/|k| = 15.36 units of x.
This is exactly what a GDC's exponential regression does (a straight line fitted to ln y), applied to y − c.
Step 7 · Residuals and goodness of fit (in the original units)
For each point, the residual is the observed value minus the model's value: e = y − ŷ.
| i | x | y (data) | ŷ (model) | y − ŷ | (y − ŷ)² |
|---|---|---|---|---|---|
| 1 | 0 | 85 | 85.13 | −0.127 | 0.0161 |
| 2 | 5 | 72.4 | 72.18 | 0.224 | 0.0504 |
| 3 | 10 | 61.5 | 61.84 | −0.340 | 0.116 |
| 4 | 15 | 53.9 | 53.59 | 0.308 | 0.0951 |
| 5 | 20 | 46.7 | 47.01 | −0.309 | 0.0957 |
| 6 | 25 | 42 | 41.76 | 0.244 | 0.0594 |
| 7 | 30 | 37.4 | 37.56 | −0.164 | 0.0270 |
| 8 | 35 | 34.5 | 34.22 | 0.281 | 0.0790 |
| 9 | 40 | 31.4 | 31.55 | −0.149 | 0.0222 |
| Σ | 0.5604 |
- Sum of squared residuals: SSR = Σ(y − ŷ)² = 0.5604
- Total sum of squares about the mean ȳ = 51.64: SST = Σ(y − ȳ)² = 2670
- Coefficient of determination: R² = 1 − SSR/SST = 1 − 0.5604/2670 = 0.9998
- Root-mean-square error: RMSE = √(SSR/n) = 0.250 — a typical size of a residual, in the units of y.
r above measures how straight the transformed data are; R² here measures the fit to the original y values. They are different numbers — quote the one you mean.
Step 8 · Refine: least squares on the original data
The ln method minimises squared errors in ln(y − c), not in y, so it gives the small y values relatively more weight. Minimising SSR in the original units (iteratively, starting from the values above, and letting c vary too) gives:
y = 64.12e−0.04499x + 20.95
Desmos does this by default for y₁ ~ a e^(k x₁) + c; turn on its “log mode” to get the GDC's answer instead. Say which one you used.
Step 9 · Residuals of the refined model
For each point, the residual is the observed value minus the model's value: e = y − ŷ.
| i | x | y (data) | ŷ (model) | y − ŷ | (y − ŷ)² |
|---|---|---|---|---|---|
| 1 | 0 | 85 | 85.07 | −0.0668 | 0.00446 |
| 2 | 5 | 72.4 | 72.15 | 0.249 | 0.0621 |
| 3 | 10 | 61.5 | 61.84 | −0.337 | 0.113 |
| 4 | 15 | 53.9 | 53.60 | 0.300 | 0.0900 |
| 5 | 20 | 46.7 | 47.02 | −0.323 | 0.104 |
| 6 | 25 | 42 | 41.77 | 0.230 | 0.0528 |
| 7 | 30 | 37.4 | 37.58 | −0.176 | 0.0309 |
| 8 | 35 | 34.5 | 34.23 | 0.274 | 0.0750 |
| 9 | 40 | 31.4 | 31.55 | −0.151 | 0.0229 |
| Σ | 0.5555 |
- Sum of squared residuals: SSR = Σ(y − ŷ)² = 0.5555
- Total sum of squares about the mean ȳ = 51.64: SST = Σ(y − ȳ)² = 2670
- Coefficient of determination: R² = 1 − SSR/SST = 1 − 0.5555/2670 = 0.9998
- Root-mean-square error: RMSE = √(SSR/n) = 0.248 — a typical size of a residual, in the units of y.
What the model tells you
- The first model decays towards 0 °C — impossible in a 21 °C room. Its SSR is 36.7 against 0.560 for the second, and after 3 hours it would predict 0.92 °C, far colder than the room. This is exactly the kind of limitation to point out.
- With c = 21: k ≈ −0.0451 per minute, so the temperature difference from the room halves about every 15.4 minutes (ln 2/|k|).
- Letting c vary gives a fitted room temperature of 20.9 °C, very close to the measured 21 °C — independent support for the model, and a strong point of reflection.
The same on a GDC
Enter and plot the data first, then fit. The key sequences are for current operating systems; menus differ slightly between versions.
TI-84 Plus CE
- Data: [stat] → 1: Edit… Type the x values in L1 and the y values in L2. To plot: [2nd] [y=] (STAT PLOT) → Plot1: On, Type: scatter, Xlist: L1, Ylist: L2, then [zoom] → 9: ZoomStat.
- y = a·bˣ with asymptote 0: [stat] → CALC → 0: ExpReg, Xlist L1, Ylist L2, Store RegEQ Y1. It gives a and b; k = ln b.
- With c = 21: in L3's header type L2 − 21, then ExpReg with L1, L3; the model is y = a·bˣ + 21.
- To see the linearisation: L4 = ln(L3) and LinReg(ax+b) L1, L4 — gradient k, intercept ln a.
TI-Nspire CX
- Data: Add a Lists & Spreadsheet page; name column A xs and column B ys and type the data. Add a Data & Statistics page (or a Graphs page with menu → Graph Entry/Edit → Scatter Plot) and choose xs and ys.
- Exponential Regression (y = a·bˣ) with X List xs and Y List ys, or Y List ys − 21 via a new column.
- k = ln(b); a·bˣ = a·e^(kx).
Casio fx-CG50
- Data: [MENU] → Statistics. Type the x values in List 1 and the y values in List 2. To plot: [F1] (GRAPH) → [F6] (SET): Graph Type Scatter, XList List1, YList List2; [EXIT], then [F1] (GRAPH1).
- Statistics → [F2] (CALC) → [F3] (REG) → [F6] (▷) → [F2] (Exp) → [F1] (ae^bx) or [F2] (ab^x).
- For c = 21, first fill List 3 with List 2 − 21 and use XList List1, YList List3 in [F6] (SET).
In Desmos
Free at desmos.com/calculator. In a regression, ~ means “fit this model”; subscripts are typed with an underscore (x_1 shows as x₁).
- Data in x₁, y₁.
- Asymptote from context:
y_1 ~ a e^(k x_1) + 21. - Let c vary too:
y_1 ~ a e^(k x_1) + c. - Desmos minimises SSR in the original units; turn on “log mode” (in the regression's settings) to reproduce the GDC's ln-based answer. Say which you used.
More on technology in the IA: using Desmos, GeoGebra and Excel.
How to write it up in your IA
- Justify the exponential form from the context (Newton's law of cooling says the rate of cooling is proportional to the temperature difference).
- Say where the asymptote comes from and show the transformed table ln(y − c).
- Show the straight-line fit and the back-substitution: k from the gradient, a from e^(intercept).
- Interpret k (or the half-life / doubling time) in context.
- Compare the asymptote-0 model with the asymptote-21 model using SSR and the residuals.
- Explain the difference between the ln-based fit and direct least squares, and which you report.
These are the points to cover, not sentences to copy. Write every explanation in your own words, about your own data.
Common mistakes
- Using ExpReg on data that approach a non-zero value (room temperature, a maximum).
- Forgetting that b = e^k: a·bˣ and a·e^(kx) are the same model with different letters.
- Taking ln of y ≤ c (undefined) — check every value of y − c is positive.
- Comparing r of the ln-transformed data with R² of another model's original data.
Is this good enough for Criterion E?
- The exponential form and the asymptote are justified from the context.
- The linearisation is shown with a table and a straight-line graph.
- Parameters are found by back-substitution and interpreted (rate, half-life or doubling time).
- An alternative (asymptote 0, or a different model) is compared with SSR or residuals.
- The difference between log-linear and direct least squares is acknowledged.
- HL: the model is derived from a differential equation (dT/dt = k(T − c)) and the fitted k is interpreted in it.
SL or HL? Criterion E asks for mathematics that fits your course. At SL, fitting with technology is fine when you explain the method and justify every choice. At HL, show more of the mathematics yourself — the last item in the list is an example. See Criterion E and Criterion D.
Frequently asked questions
Why do my GDC and Desmos give different exponential models?
A GDC fits a straight line to ln y, which minimises errors in ln y. Desmos by default minimises errors in y itself. Both are valid; they weight the points differently. Turn on Desmos's log mode to match the GDC, and say which you used.
Free: the IA checklist an examiner uses
Every check for Criteria A–E in a 4-page PDF, the mistakes that cost the most marks and a self-assessment grid. We'll email it with a short IA tip every few days, timed to your deadline if you give it. Free — no account, no payment.
While you wait for the email: read the full modelling workflow →