Methodology All editions ENNL

After The Forecast

The weather forecast, reviewed after the weather.

Published monthly by VeriSky



Methodology

How the scoring works, and what it cannot tell you

A verification report is only worth reading if you can see how it was made. This page is the full method, including the parts that limit what the numbers mean.

What is compared against what

For every hour, we record what each model predicted for that hour, at each distance from one to seven days ahead, and keep the prediction until the hour arrives. An After The Forecast ranking is published only when that prediction can be compared with an observation independent of every model in the table.

Edition No. 1 uses temperature and wind reports from Schiphol and Rotterdam, supplied by the NOAA Aviation Weather Center. A report is snapped to the nearest whole hour only when it falls within 20 minutes. A station-day qualifies only when all 24 hours have a usable temperature and wind value.

VeriSky's wider scorer also banks results against Open-Meteo's best-match analysis where no station record exists, and for rain. That analysis is itself a weather-model product. Over the Netherlands it resolves to KNMI Harmonie, so it is not an independent source for a public model ranking. Those analysis-backed results are not published by After The Forecast.

The two published numbers

Temperature is the root mean square error in degrees, then placed on a 0 to 100 scale where an error of 6.5 degrees or more scores zero and a perfect forecast scores 100. The raw error in degrees is printed beside the score, because the score is a convenience and the error is the measurement.

Wind is scored the same way, against a scale that widens when the wind itself is strong, so a 10 km/h miss in a gale is not treated as the same failure as a 10 km/h miss in still air.

Why rain is not ranked yet

METAR precipitation reports are too sparse and coarse for the hourly wet-or-dry score. The scorer's current fallback is the best-match model analysis, which would let one ranked model supply the truth. Radar collection is running, but it was not the source of the July daily accumulators. Rain remains absent until a complete month can be rebuilt from the independent radar record and pass the same coverage checks.

How the common sample is built

First, only complete station-days are eligible. A model must serve at least 95 percent of those hours to become a candidate. The final table is then reduced to the exact intersection of station-days served completely by every candidate model. That gives every row the same places, dates and hours rather than merely similar totals.

Headline tables compare one lead time, usually one day ahead. Pooling several distances would reward short-horizon models that never face the harder later days. Lead-time charts keep the distances separate and stop a line when a point falls below 95 percent of the headline sample. Exact hours are printed with the chart data.

Which models qualify

Models that appeared partway through the month or missed too much of the complete station sample are scored internally but not ranked. A renderer allowlist gives every published model an explicit name, origin and accessible colour; an unknown model stops the build.

Some national weather services publish a seamless product that serves their own high resolution model where it reaches and another centre's model elsewhere. Outside its own domain such a product is not that country's forecast. Where we can show it is identical to the model it is standing in for, hour for hour, it is dropped rather than listed as a second entry under a national name.

Commercial forecast providers are scored but are not published in these rankings. Several of their licence terms forbid ranking or comparing their output, and rather than publish a partial table that includes some and not others, we publish none of them.

Our own forecast

VeriSky publishes a blended forecast that weights models by how they have been scoring lately. It is held out of After The Forecast's provider standings. The publication can discuss the blend, but it does not let the publisher's own product compete for the headline.

What the numbers cannot tell you

Locations are not the whole country. Edition No. 1 uses two airport stations. They are not a uniform grid and do not represent inland minima, coastal waters or every Dutch weather regime. A model strong at those airports may be weaker elsewhere.

Complete-day filtering narrows the month. Missing station reports and incomplete model coverage remove whole station-days. Each edition publishes the retained station, day and hour counts so the size and geography of the evidence remain visible.

One month is one month. A model that wins July has won July. Weather regimes favour different models, and a single edition is not a verdict on a model's quality.

Scores are relative to a scale we chose. The 0 to 100 conversion makes models comparable to each other, not to some absolute standard of forecasting. The underlying error is always printed so you can ignore our scale if you prefer.

Reproducibility and corrections

Every edition is generated from poolable daily accumulators in the scoring database, not from a rolling display table. Station coverage is derived from the observation rows, not a hand-set label. The dataset each page was rendered from is published beside it as data.json.

Editions are never silently rewritten. When something is wrong we fix it, mark the change on the page, and list it at corrections. Reports are published only after a month is complete, never for a month still in progress.