German Grid Forecasts: Where the Uncertainty Bands Fail
A 13-month audit of German day-ahead forecasts against what happened: systematic midday bias, uncertainty bands that fail out-of-sample (including conformal ones), and what that costs: ≈€45,500 per 100 MW of avoidable imbalance exposure at reBAP, concentrated in exactly the hours where failing is expensive.
Version history
- v0.1 · 2026-08-17 · working draft for review
- v0.1.1 · 2026-09-01 · correction: an earlier version said Germany does not publish imbalance prices openly. It does. reBAP is published quarter-hourly by the four TSOs on netztransparenz.de. The Dutch leg carries the priced findings because it is the leg that was run, which is what the section now says.
- v1.0 · 2026-09-02 · first public release. Content unchanged from 0.1.1; the version marks the paper as issued rather than circulating for review.
- v1.1 · 2026-09-02 · the monetised leg is re-priced on German reBAP, published quarter-hourly by the four TSOs. The Dutch figures it previously carried are withdrawn: that series pair fails a forecast-integrity check (correlation 0.62, slope 0.61, forecast below actual at all 96 slots), so a euro figure scaled off mean load could not stand on it. Coverage results are unchanged.
We audit every quarter-hour of German (DE-LU) day-ahead system forecasts from July 2025 through July 2026 (load, solar, onshore and offshore wind) against realized outcomes from the ENTSO-E Transparency Platform. Three results. First, the point forecasts carry conditional bias that vanishes in aggregate statistics: load is over-forecast by roughly 1.9-2.0 GW every midday. Second, the two standard ways of wrapping intervals around a point forecast, parametric Gaussian bands and rolling split-conformal quantiles, both undercover out-of-sample (83-89% at nominal 90%), and collapse regionally: solar's nominal-80% band covered 56-63% through spring 2026. Distribution shift breaks the exchangeability that conformal's guarantee rests on. Third, adaptive conformal inference restores marginal coverage at an honestly stated width cost, and reduces without removing the tendency to miss where a miss is expensive. Priced at German reBAP, trusting the Gaussian "90%" band carried roughly EUR 90,000 of un-reserved adverse exposure per 100 MW over the period, versus EUR 44,500 under the adaptive band, with misses concentrated in the intervals where imbalance was most expensive: the mean adverse spread was EUR 24.6/MWh overall and EUR 34.8/MWh on the intervals the band missed.
- Calibration background/calibration · v1.0 · tutorial
Every energy desk in Europe consumes the TSO day-ahead forecasts, and almost nobody asks the question this analysis answers: when those forecasts imply a range of outcomes, is the range honest? We audit thirteen months of German (DE-LU) day-ahead forecasts against what actually happened (every quarter-hour from July 2025 through July 2026) for system load, solar, onshore wind and offshore wind: 38,016 quarter-hours per series, worst-case missingness 0.05%, all from the ENTSO-E Transparency Platform.
The findings run in increasing order of consequence: a conditional bias that averages conceal (§ 02); the failure, out-of-sample, of both standard ways to build uncertainty bands, including the naive conformal one (§ 03); and the adaptive construction that keeps the average promise, at a stated price and without closing the gap where it is expensive (§ 04). § 05 records what this audit can and cannot show.
The data, and the rules of evidence
Public forecasts, public outcomes, one strict discipline: every number scored strictly out-of-sample.
The Transparency Platform publishes each TSO’s day-ahead point forecasts alongside realized values. That pairing makes an audit possible without anyone’s permission: the forecast was committed before delivery, the outcome is measured after it, and the difference (the error e = actual − forecast) accumulates into a track record nobody curates.
Two rules govern everything below. First, any quantity fitted from history (a mean, a standard deviation, a quantile) uses a trailing 90-day window that ends the day before the observation it is scored on. No method ever sees its own exam. Second, every construction is evaluated per quarter-hour delivery slot, because grid errors have a schedule: a method that is right on average and wrong every noon should be caught, not excused.
The bias clock
Germany's load forecast overshoots by two power stations' worth. Every day, at the same hours.
The aggregate numbers look respectable: load bias −385 MW on a ~55 GW system, solar −56 MW, offshore wind +176 MW. The clock tells another story.
Mean forecast error by delivery hour (Europe/Berlin), 13 months of quarter-hours. Bars below the zero line mean the day-ahead forecast was too high at that hour. The overall mean (dashed) is what a summary statistic reports. The clock is what dispatch experiences.
Load is over-forecast by 1.8–2.0 GW at hours 11–14: roughly two large power stations of demand that is predicted daily and daily fails to materialize, five times the size of the aggregate bias. The likely mechanism is well known (behind-the-meter solar eroding metered load faster than the forecast model learns); the size is something you only learn by measuring. Offshore wind leans the other way, under-forecast by ~240 MW through the night hours; solar’s midday under-forecast peaks around −420 MW.
A summary bias statistic averages over the clock. A desk that trades the midday blocks does not. Any audit, internal or vendor, that reports a single bias number per series has not yet said anything about the hours where the money is.
Three ways to promise 90%
The parametric habit fails. So does the naive conformal alternative. Admitting that is the point.
The TSO publishes points, not ranges. Any desk wanting a range constructs it, and the standard construction is parametric: fit mean and standard deviation to recent errors, quote . The textbook distribution-free alternative replaces the Gaussian fit with trailing empirical quantiles: a rolling split-conformal band. We ran both, identically windowed, strictly out-of-sample; and a third method introduced in § 04.
A desk with a quantitative team will say it does not build bands this way, and it is right. It runs quantile regression on weather features, or it buys a probabilistic forecast with bands already attached. That objection does not weaken the audit; it describes the normal case for it. What is under test is not how a band was built but whether the number printed on it is true. The procedure here scores whatever band you already quote, from whatever construction, against the outcomes you already have: it counts how often the band contained the answer, splits that count by regime, and prices the misses. Nothing in it needs to know the model. The two constructions audited below are in the paper because they are public and reproducible, not because they are the state of the art, and a better-built band is a more interesting subject for the same test rather than an exemption from it.
Coverage each method actually delivered at the selected promise, per series. All three use the same trailing 90-day windows per quarter-hour slot, scored strictly one day ahead. The madder segment is the gap between promise and delivery. Interval width is the honest price of closing it: shown in the readout.
At nominal 90%, the Gaussian band delivered 84.8–88.9% depending on series; rolling conformal delivered 83.2–86.6%. Neither is a 90% band. The aggregate conceals worse: solar’s nominal-80% construction covered 56–63% through spring 2026, precisely when new solar capacity made the error distribution grow faster than any trailing window could learn.
Split-conformal’s celebrated finite-sample guarantee is conditional on exchangeability: tomorrow’s error drawn from the same distribution as the window’s. A growing solar fleet is precisely a violation of that condition. Applied statically under distribution shift, conformal undercovers just like the Gaussian it replaced. A warranty is only as good as its conditions.
The construction that keeps its promise
Adaptive conformal inference tracks its own misses and repairs them. Marginal coverage returns to nominal; the expensive hours stay harder.
Adaptive conformal inference (Gibbs & Candès, 2021) treats the miss rate as a control problem: after every observation, the target miss rate is nudged,
so the very next interval widens after a miss and tightens after a long run of hits. The step size γ is a plain tuning constant (larger reacts faster and wobbles more), and it was fixed at γ = 0.02 before any evaluation, not tuned to the results. Same trailing windows, same no-lookahead discipline, same data.
The result, visible in Fig. 02 for every series and level: achieved coverage lands at 89.0–89.7% against a 90% promise, 79.3–79.9% against 80%, 93.1–94.6% against 95%. The price is stated rather than hidden, and stated in the tail, not just the mean, because an average width in an article about the sins of averaging would be an irony. The adaptive load band at 90% averages 8.8 GW against the Gaussian’s 7.5 GW; at the 95th percentile of widths it quotes 14.2 GW against the Gaussian’s 11.1 GW, and its single widest interval reached 19.3 GW (onshore wind’s worst: 25.8 GW). Those wide moments are the method telling you it recently missed and does not yet trust the window. That is information, not noise. The Gaussian’s narrower quotes at those same moments were never actually 90% bands. Calibration does not make uncertainty larger; it stops you understating it.
Where the aggregate table still flatters everyone is regime. The month-by-month cut is the figure this article’s own Recipe demands:
The same three constructions, scored month by month. Aggregate tables average this away; the promise is only as good as its worst regime. The shaded band marks spring 2026, when German solar additions shifted the error distribution faster than a trailing window learns.
The month-by-month cut also keeps us honest about what adaptivity is. When the solar shift hits in March 2026, all three methods get hurt: ACI’s worst month is 65% against an 80% promise, because a reactive method cannot see a regime break coming. The difference is what happens next: ACI walks back to 74% in April, 79% in May, 82% by June, while the static constructions sit in the mid-50s to low-60s for four consecutive months, still quoting “80%” the whole time. Adaptivity is not immunity; it is recovery measured in weeks instead of quarters, and a stated miss you can see happening, instead of a silent one.
- Pair each committed forecast with its realized outcome; never re-fit after the fact.
- Score whatever band you quote (parametric or otherwise) strictly out-of-sample, per delivery slot.
- Report coverage by regime (hour, season, fleet state), not only in aggregate.
- If coverage drifts from the promise, adapt the miss rate, not the story.
What this audit can and cannot show
TSO forecasts are the public benchmark, not your stack. The method transfers, the numbers will not.
What do these percentage points cost? German imbalance answers, because reBAP is published quarter-hourly by the four TSOs on netztransparenz.de. Pricing every escape from the trailing Gaussian “90%” band at that quarter-hour’s adverse imbalance-versus-day-ahead spread (favourable spreads clipped to zero: risk side only), scaled to a portfolio with 100 MW mean load whose errors move proportionally to the system’s (an assumption, stated): €90,000 of un-reserved adverse exposure over the 13 months, against €44,500 under the adaptive band: a ≈€45,500 difference per 100 MW. The stated band missed 15.2% of quarter-hours. And the property the averages hide: the mean adverse spread over all intervals was €24.6/MWh, while on the intervals the Gaussian band missed it was €34.8/MWh. The band fails disproportionately in exactly the intervals where failing is expensive. Miscalibration and market stress share causes; that correlation is the audit’s sharpest commercial fact.
Until v1.1 this figure came from a Dutch run, because NL imbalance prices were pulled first. That leg has been withdrawn. The two ENTSO-E series behind it are not a forecast and its outcome: the published NL load forecast sits below realised load at every one of the 96 quarter-hour slots, correlating at 0.62 on a slope of 0.61, against 0.96 and 1.02 for the German pair. The coverage arithmetic survived that, because a trailing per-slot window absorbs a stable offset, but a euro figure scaled off mean load does not. The specimen audit now runs a series-integrity check before scoring anything, and keeps the Dutch run as the case where it fires.
Beyond that, this analysis can show whether a published point forecast is conditionally biased, and whether a band constructed around it holds its promised frequency. Those are checkable against outcomes. It cannot show whether your forecasts share these failure modes: portfolio forecasts differ from system forecasts, and a desk’s decision layer consumes uncertainty differently than a TSO’s.
One consequence deserves naming, because it is where this method has teeth. If you buy forecasts from a vendor who quotes confidence bands, everything above applies to them exactly as it applies to the TSO. Coverage is checkable by the customer, from data the customer already holds (committed forecasts and realized outcomes) without the vendor’s cooperation or their model. A vendor whose “95%” bands blow through weekly in a windy quarter is not suffering “unprecedented weather”; they are failing an audit nobody has run. Recipe 4.1 doubles as the four questions to put in your next RFP or quarterly vendor review.
The audit methodology (the pairing discipline, per-slot scoring, regime-conditional coverage, the adaptive repair) transfers unchanged to any stack that can export twelve months of forecasts and outcomes.
Forecast Calibration Audit: your committed forecasts and outcomes (or your vendor’s), audited exactly as above. Input: one CSV per series (timestamp, forecast, actual, stated band if any); 12+ months. You get: the bias clock and coverage ladder for your portfolio; coverage by regime and by season; the width-tail accounting; a monetized shortfall estimate against your market’s prices; an adaptive recalibration demonstration on your own data; a 45-minute readout. Terms: fixed price by portfolio size, two weeks from data delivery; runs on your infrastructure if preferred; your data is deleted on delivery and never used elsewhere; standard mutual NDA before anything moves. A scoped single-series Vendor Band Audit is available as a smaller first step; a specimen of that deliverable, run end to end on a public series, shows exactly what arrives. If your bands are honest, the audit proves that. A marketable result too.
Matthew Tanti · Kwantil, Malta · contact via kwantil.com
@article{tanti2026,
author = {Matthew Tanti},
title = {German Grid Forecasts: Where the Uncertainty Bands Fail},
journal = {Kwantil},
year = {2026},
number = {KW-2026-01},
version = {1.1},
url = {https://kwantil.com/papers/german-grid-forecast-calibration/}
}Matthew Tanti (2026). German Grid Forecasts: Where the Uncertainty Bands Fail. Kwantil, KW-2026-01, v1.1.