The main page shows one city at a time, which is the right way to follow a single model through a year and the wrong way to see what the six have in common. This is the other view. Every panel below carries all six, so the differences between them are the subject rather than a thing you hold in memory across six clicks.
One panel per city, same scale on every one, so a shallow curve and a collapsing one are told apart by shape rather than by reading the axis. Above the line the model beats a rule of thumb, which is what this hour usually looked like over the previous 30 days. Below it, it loses to one.
Two readings per city, because the first misleads on its own. Across the replay compares the median error of what was served against the first model held frozen, which is an unpaired comparison: where a city promotes nothing until late, most windows compare the first model against itself and the answer collapses toward zero. Week by week holds the window fixed and compares the two models inside it.
Each city here gets its own model, retrained on its own schedule. The cheaper arrangement is one model over all six, so every city's benchmark carries it: one Ridge, trained once on the six training windows, never retrained.