DEMONSTRATION DATA
Historical Hindcast / Backtesting
Select a past flood episode. PRAVAH is re-run as if only information available up to shortly before that episode began were known: every read is clipped to that cutoff (see DataProvider.as_of), and the resulting forecasts are scored against what the record shows actually happened afterwards. This is a blocked, chronological evaluation walking forward through one real episode, never a random train/test split. The live/ongoing episode is not offered here because it hasn't finished yet.

Onset is the episode's real recorded start time.

Baseline: persistence + routing/rainfall physics only, every component traceable. ML residual hybrid: baseline plus a learned correction, usually more accurate but less explainable.

Past Backtest Runs

Loading…

Glossary

MAE (Mean Absolute Error)
Average of |predicted − observed| across every scored forecast, in metres. Treats an over-prediction and an under-prediction of the same size as equally bad.
RMSE (Root Mean Squared Error)
Same idea as MAE but squares each error before averaging, so a few large misses raise it more than many small ones. RMSE ≥ MAE always; a big gap between them means occasional large misses, not just steady small ones.
Bias
Average of (predicted − observed), signed. Positive: the model runs high on average (over-predicts). Negative: it runs low (under-predicts). Near zero doesn't mean accurate, a model that's often 1m high and often 1m low averages to zero bias with a large MAE.
Horizon (+6h, +12h, …)
How far ahead of a forecast's own issuance time the predicted value is for. +72h is a harder prediction than +6h and is expected to score worse.
Issuance
One discrete forecast run made at a specific point in time during the walk-forward evaluation. A backtest scores every issuance PRAVAH would have made while the real episode was unfolding, not just one.
Peak timing error
Difference between when the model's forecast said the episode's peak water level would occur and when the record shows it actually occurred, in hours.
Danger-threshold crossing timing error
Same idea, but for when the level was forecast to cross the station's danger threshold versus when it actually did. "No threshold crossing in window" means the real level never reached danger during the scored window, so there's nothing to time against.
P10 / P50 / P90
Percentiles of the forecast's uncertainty range. P50 is the central (median) estimate; there's a modelled 80% chance the real value falls between P10 and P90. A wider P10–P90 band means less confidence, not a wider error necessarily.
Onset
The real recorded time the historical episode began, used as the walk-forward starting point.
Blocked / walk-forward evaluation
Every forecast in a backtest only sees data available up to its own issuance time (DataProvider.as_of), scored against what actually happened afterwards, in chronological order. Never a random train/test split, which would let a model "see the future" of the same episode it's being scored on.