Predicting transit delays — and beating the operators' own forecasts.
Independently trained models forecast arrival delays up to nine stops ahead across Swedish transit networks, each graded head-to-head against that operator's published real-time forecast on identical events and the same settled truth.
Why these numbers are honest
Each “beats” claim has to survive two independent significance tests that make different assumptions about the data. Where the evidence isn't there — often at the longest horizons, where data thins — the board says so.
Same instant.The operator's forecast is captured at the exact moment the model makes its own, by point-in-time lookup — never at a later, more informed time.
Settled truth.Grading waits until the real delay has settled, so the operator isn't scored against its own un-updated guess.
Operator-covered only.Every average uses only events where the operator also published a forecast; coverage is shown per horizon.
Two tests must agree.A horizon reads “beats” only when a Diebold–Mariano test (Newey–West HAC) and a trip-clustered bootstrap both place the whole 95% interval below zero, above minimum event and trip counts.