How the scoring works, and what it cannot tell you
A verification report is only worth reading if you can see how it was made. This page is the full method, including the parts that limit what the numbers mean.
What is compared against what
For every hour, we record what each model predicted for that hour, at each distance from one to seven days ahead, and keep the prediction until the hour arrives. An After The Forecast ranking is published only when that prediction can be compared with an observation independent of every model in the table.
Edition No. 1 uses temperature and wind reports from Schiphol and Rotterdam, supplied by the NOAA Aviation Weather Center, and reads its event day from De Bilt, KNMI's reference station, which reports every ten minutes through the KNMI Data Platform. A report is snapped to the nearest whole hour only when it falls within 20 minutes. A station-day qualifies only when all 24 hours have a usable temperature and wind value.
VeriSky's wider scorer also banks results against Open-Meteo's best-match analysis where no station record exists. That analysis is itself a weather-model product. Over the Netherlands it resolves to KNMI Harmonie, so it is not an independent source for a public model ranking. Those analysis-backed results are not published by After The Forecast, and no metric on a report page is ever taken from them.
How temperature and wind are scored
The score uses VeriSky's Accuracy + Extremes model. Temperature accuracy is the share of hours forecast within 2 degrees of the observation. Wind accuracy is the share within 5 km/h, widened to 20 percent of the observed speed when that is larger. This makes the score a direct account of how often a forecast was usefully close.
The extremes component checks whether a model caught unusually warm, cold and windy hours without repeatedly calling extremes that did not happen. When the sample contains at least eight extreme and eight non-extreme hours, this component contributes 30 percent and accuracy contributes 70 percent. On a quieter sample, accuracy carries the whole score. Root mean square error remains printed beside each score as the typical error in degrees or km/h; it is no longer converted into the score itself.
Rain is ranked against the station's own rain gauge, where one covers the whole period. The gauge measures what reached the ground, which is a measurement and not a model analysis, and that independence is what allows rain to be ranked at all. It reports the rainfall of the hour just past, and the archive stores a model's rainfall the same way, so both sides mean the same hour by the same name.
De Bilt supplies the KNMI R1H reading taken on each whole hour. RAF Lakenheath supplies a scheduled METAR at 55 minutes past the hour; its P group measures the preceding hour and is assigned to the whole hour ending five minutes later. At Lakenheath, a scheduled report without a P group is dry, and so is a P0000 group, which reports a trace too small to measure. If the scheduled report itself is missing, the hour is unknown and is left out for every model. The trace rule applies from edition No. 6; No. 4 counted a trace as a wet hour, as the corrections page records.
An hour counts as wet on either side at a tenth of a millimetre, the smallest amount the archive stores. Each side is credited within an hour of the other, so a band forecast an hour early is a hit rather than a false alarm plus a miss. An hour with no usable reading is judged by nobody rather than counted as dry, and the hours left out are named under the figure.
Rain is judged at every notice from one to seven days ahead, on the same days at every notice: a day enters only when the gauge could be read and every run before it is in the archive, so a column changes with distance and never with the sample. Beside each score the grid prints the rain that run forecast over those hours. That is also how the collected commercial forecasts take part: their hourly figures follow a different convention, so they carry an amount and never an hourly score. There is no row for how far ahead a model first had the rain. That question suits a single hot day, not showers; editions No. 3 and No. 4, which judged a whole wet week by a factor of two, are kept as they were published.
In the hourly tables a model takes part in a day only when its run covers all 24 hours of it. A model with a short horizon can therefore drop out of a column, but it can never shrink the sample the other models are scored on. Rain at a single point remains the hardest question the publication asks: an hour of drizzle is smaller than a model grid box can resolve, and no rain figure here is a verdict on a whole country.
The event
Each edition opens with one event: the hottest complete station-day in the period, or a wet spell named by the range of days it covers. A hot day is not chosen by hand, it falls out of the station record, and the runner-up is printed in the method note so the choice can be checked. A spell is an editorial choice, so the method note names it, along with the wettest day inside it and the wettest day of the whole period when those differ. The public models are replayed from the archived 00 UTC run of each preceding day, so every one of them is asked the same question at the same moment. Commercial services have no archive to replay and are read instead from the forecast we collected on the day, which the edition states beside their columns.
The grid keeps reaching further back for as long as any model still had the day inside the band, and ends on the first notice at which none of them did. That last row is the point of the section: it is where the event stops existing in anybody's forecast.
Reading an event by lead time asks a different question from an average. A model can be the most accurate over six weeks and still lose the one day that mattered, and a single lucky run is not the same as having seen the day coming. The heat grid therefore credits a model from the longest unbroken run of calls inside the band, counted back from one day ahead.
How the common sample is built
First, only complete station-days are eligible. A model must serve at least 95 percent of those hours to become a candidate. The final table is then reduced to the exact intersection of station-days served completely by every candidate model. That gives every row the same places, dates and hours rather than merely similar totals.
Headline tables compare one lead time, usually one day ahead. Pooling several distances would reward short-horizon models that never face the harder later days. Lead-time charts keep the distances separate and stop a line when a point falls below 95 percent of the headline sample. Exact hours are printed with the chart data.
Which models qualify
A report page ranks a short, fixed list: the national model of the country it covers, one model from each of the four centres most European forecasts are built on, and ECMWF's machine-learned model, which answers the same question without solving the physics. That shortlist is fixed by institution before any score is read, so it can never be a selection made after seeing who won. Every other open model scored on the same hours is published in the dataset beside the page, so the shortlist can be checked against everything it leaves out. A renderer allowlist gives every model an explicit name, origin and accessible colour; an unknown model stops the build.
Some national weather services publish a seamless product that serves their own high-resolution model where it reaches and another centre's model beyond it. Past that handover the series is not that country's forecast. Each edition measures the national model's own reach from the standalone member's coverage, prints nothing under the national name beyond it, and says so on the page. Where the archive proves the handover by repeating another model's accumulators exactly, that is reported too.
Commercial forecast providers are scored but are not published in these rankings. Several of their licence terms forbid ranking or comparing their output, and rather than publish a partial table that includes some and not others, we publish none of them.
Our own forecast
VeriSky publishes a blended forecast that weights models by how they have been scoring lately. It is held out of After The Forecast's provider standings. The publication can discuss the blend, but it does not let the publisher's own product compete for the headline.
What the numbers cannot tell you
Locations are not the whole country. Edition No. 1 uses two airport stations. They are not a uniform grid and do not represent inland minima, coastal waters or every Dutch weather regime. A model strong at those airports may be weaker elsewhere.
Complete-day filtering narrows the period. Missing station reports and incomplete model coverage remove whole station-days. Each edition publishes the retained station, day and hour counts so the size and geography of the evidence remain visible.
One period is one period. A model that wins these weeks has won these weeks. Weather regimes favour different models, and a single edition is not a verdict on a model's quality. This is doubly true of the two event sections, which are one day each.
Scores are relative to a scale we chose. The 0 to 100 conversion makes models comparable to each other, not to some absolute standard of forecasting. The underlying error is always printed so you can ignore our scale if you prefer.
Reproducibility and corrections
Every edition is generated from poolable daily accumulators in the scoring database, not from
a rolling display table. Station coverage is derived from the observation rows, not a hand-set
label. The dataset each page was rendered from is published beside it as
data.json.
Editions are never silently rewritten. When something is wrong we fix it, mark the change on the page, and list it at corrections. A report covers a closed range of whole days and is published only once every day in it has been observed and scored, never while the period is still running.