Rolling baselines, explained

A baseline is a written-down answer to “what does normal look like?” A rolling baseline keeps rewriting that answer as new data arrives. How it is computed decides what it can see, and the differences are easiest to understand with numbers.

Explainer 7 min read

A baseline is a reference for normal behaviour of a metric: a typical level and a typical amount of variation around it. New readings are compared with it. A rolling baseline is recomputed over a moving window of recent history, such as the last eight weeks, so that it follows the business as it grows, shrinks or changes shape.

Nearly every useful alert on business data is a baseline underneath. The details below are what separate one that works from one that cries wolf or sleeps through problems.

Why rolling

A fixed baseline, computed once, goes stale. A shop that has grown 30% since the baseline was set will see every ordinary day flagged as unusually high. A rolling window avoids that by always describing recent normal rather than historical normal.

The trade-off is the window length:

  • Short windows (a week or two) react quickly to real change, but are noisy, and they absorb problems fast: a drop that lasts ten days becomes the new normal before anyone has acted on it.
  • Long windows (several months) are stable, but slow. After a genuine, permanent change, such as a price rise, they keep flagging the new level for weeks.

For daily business metrics, six to twelve weeks is a common, sensible range. Pair it with a separate test for level shifts, so that a genuine change is reported once, clearly, instead of slowly leaking into the baseline.

Mean or median: a worked example

The obvious way to compute a baseline is the average and the standard deviation. It has a serious flaw: both are pulled hard by the unusual values the baseline exists to catch. Take seven days of a steady metric with one spike:

StatisticMean and standard deviationMedian and robust spread
Data100, 102, 98, 101, 99, 100, 300100, 102, 98, 101, 99, 100, 300
Typical level128.6100
Typical spread75.61.48
Next reading of 150 scores0.333.7
Verdict on 150NormalVery unusual
The robust spread is the median absolute deviation, 1, multiplied by 1.4826.

One spike made the mean-based baseline think normal is 128.6 and that swings of 75 are ordinary. So the next jump of 50% scores as nothing. The median ignores the spike entirely: the level is 100, days typically wander by about 1.5, and 150 is correctly far outside that.

A practical wrinkle: if more than half the values are identical, as with a metric that is usually zero, the median absolute deviation is zero and every movement looks infinite. A robust baseline falls back to another resistant measure, such as the interquartile range, rather than to the standard deviation, which would bring the outlier problem straight back.

Seasonal baselines

Most business metrics have a weekly rhythm. A baseline that treats every day alike will flag busy Saturdays and miss bad ones. The fix is a seasonal profile: learn how each weekday typically compares with the overall level, remove that pattern before measuring the spread, and add it back when judging a new reading.

Actual Normal range Outside
Example: a baseline with a weekday profile. The band rises every weekend, so the last Saturday, at a weekday’s level, is the one reading outside it.

A profile needs repeats to be trustworthy. With one week of history, “Saturday is high” is a single observation; with three or more weeks it starts to be a pattern rather than noise.

From baseline to band

The band drawn on a chart is the level plus and minus some number of spreads, often called k:

  • k = 2 flags about one reading in twenty on well-behaved data. Sensitive, and noisy.
  • k = 3 flags about one in 370. Quiet, and still quick to catch a real break.

Choosing k is choosing between interruptions and misses. There is no right answer for every metric, which is why it should be adjustable, and why it helps to replay a year of history to see what each setting would have flagged.

How much history before it judges

With five data points, the median and spread are barely estimates. A baseline should withhold a verdict until it has enough readings, commonly somewhere around two weeks of daily data, and say so plainly rather than guess. “Still learning” is an honest state; a confident verdict from four data points is not.

In short

  • A rolling baseline is typical level and spread, recomputed over a moving window.
  • Use the median and median absolute deviation, so the anomalies you want to catch cannot hide themselves.
  • Build the weekly pattern in, once there are several weeks to learn it from.
  • Pair the baseline with a level shift test, and withhold verdicts until there is enough history.

Let a gauge do the watching.

Free to start: two gauges, no card needed.