Predicting transit delays — and beating the operators' own forecasts.
Independently trained models forecast arrival delays up to five stops ahead across Swedish transit networks, each graded head-to-head against that operator's published real-time forecast on identical events and the same settled truth.
Why these numbers are honest
Each “beats” claim has to survive two independent significance tests that make different assumptions about the data. Where the evidence isn't there, the board says so.
Same instant.The operator's forecast is captured at the exact moment the model makes its own, by point-in-time lookup — never at a later, more informed time.
Settled truth.Grading waits until the real delay has settled, so the operator isn't accidentally scored against its own un-updated guess.
Operator-covered only.Every average uses only events where the operator also published a forecast; coverage is shown beside each cell.
Two tests must agree.A cell reads “beats” only when a Diebold–Mariano test (Newey–West HAC) and a trip-clustered bootstrap both place the whole 95% interval below zero — and only above minimum event and trip counts.