Models go stale. The world they were trained on moves on, and usually nobody notices until something downstream breaks. This one notices on its own, every week, with nobody watching it. The job: forecast how dirty a city's air will be (PM2.5, the soot fine enough to reach your bloodstream) a week ahead, from nothing but the weather forecast. When the model starts slipping it trains a replacement, then refuses to ship it unless it beats the one already in service. Six cities, running since 2025.
Six cities, each with its own model and its own history of retraining. Every setting is identical across all of them, so where two cities behave differently it is their air that differs and not their tuning. Pick one, or see all six side by side.
One marker per city being watched. Colour shows how far that city's weather has wandered from what its model was trained on; size shows how long it has been under watch. Click a marker or a row to switch cities, and drag to spin the globe.
Nothing below this line depends on which city is selected. These are the findings about the loop itself, and each one needs more than a single city to make. The all-cities view puts the six on shared axes.
Each city here gets its own model, retrained on its own schedule. The obvious cheaper arrangement is one model over all six at once, so the benchmark on every city carries it as a predictor: one Ridge, trained once on all six training windows, never retrained.
Every promotion on this page was decided by one exam: a week of air neither model had seen, with the challenger needing to win by more than 5%. That margin is the number the decision was made on, so it cannot also be the evidence the decision was right. That would be marking your own homework. The check is what each winner went on to do over the weeks it served, measured against the model it displaced and scored on those same weeks.
When an alarm goes off in a real city, there is no way to know whether it was right. Nobody knows what the true answer was supposed to be, so an alarm firing at random would look much the same as one that worked. So the same system is pointed at a made-up world with two dials. One changes the weather. The other changes how the weather turns into pollution. Here the right answer is known in advance, because we set it. Each alarm should react to its own dial and stay silent for the other. That is what the charts below test.
Everything above rests on one model: a straight line through eight weather measurements. Straight lines are bad at seasons, so “the air changed and a fresh model tracks it” has a duller twin — “a line cannot bend, and refitting is how you fake it”. Those predict the same charts. Telling them apart needs a model that can bend, tried as hard as the line was, on the same two cities: Delhi, where retraining pays most, and Los Angeles, where it costs.
One scheduled run, every week. Steps 3 and 4 only happen when step 2 says they should.
Why two alarms and not one: the first looks only at the weather coming in, so it can fire immediately, without waiting to find out whether the model was wrong. But “the world looks different” is not the same as “the model is failing”. Kraków shows the gap. Through the summer its weather drifts further from training than anywhere else on this page, while the model quietly gets better. So the cheap alarm watches, and only the expensive one, the one asking whether we got this wrong, can authorise spending money on a retrain.
What this page shows is wrong with it: the expensive alarm compares the model against its own error at training time, and every promotion resets that comparison. Because retrains fire in the dirty season, each new model inherits a higher bar than the one it replaced, and the bar ratchets upward — a staircase you can watch in “when it retrains, and against what bar”. Once it has ratcheted to the seasonal peak the alarm cannot fire again, whatever the model does. Kraków spends its last 30 weeks that way, serving a 210-day-old model while the trigger reports it as healthier than it has ever been.
Five fixes were built, and none of them pays. Each failure named the next thing to try, and every one was replayed across all six cities against the shipped loop, week by week, with intervals attached. A second alarm — skill against a plain daily profile, a yardstick that promoting a model cannot move — takes Kraków's longest silence from 30 weeks to 5 and helps the one city that had gone deaf, and nothing else; it is built, and still ships switched off. A longer exam does nothing measurable anywhere. Re-examining the model on a schedule bounds how stale it can get and buys no accuracy. Making the exam harder to pass is worse than leaving it alone. Letting the loop undo a promotion does no harm and proves nothing.
They fail for one reason, and it took all five to see it. The loop keeps training rivals and holding exams until one passes, so whichever model gets the job is the one that got lucky on the day. Making the exam harder does not mean fewer bad hires; it means more attempts, and a winner that got luckier still. Nothing about when the loop looks, how strictly it marks, or whether it can change its mind afterwards can fix that. Underneath all five is the same limit: a week of hourly air is not enough to tell two of these models apart. Los Angeles is the proof. Its one promotion was made on a margin that could just as easily have been zero, and re-judging the same decision a fortnight later reverses it, then reverses it back a week after that. That is a limit of the problem rather than a bug in the machinery, and finding it is what building all five was for.
It answers one question: given the weather we expect in a week, how dirty will the air be? It is fed the weather forecasts as they were published at the time, rather than the weather that turned out to happen, so it starts a step behind in the same way anyone forecasting a week out really would. It is never told how dirty the air has been recently, which would be an easy shortcut and would hide the effect being demonstrated. Keeping it simple is a choice: a simple model visibly falls apart when the world shifts, where a clever one papers over the cracks and you learn nothing. What is worth looking at is the machinery around the model. Swap in something bigger and none of that machinery changes.