Analysis·Energy ENTSO-E

German Grid Forecasts: Where the Uncertainty Bands Fail

A 13-month audit of German day-ahead forecasts against what happened: systematic midday bias, uncertainty bands that fail out-of-sample (including conformal ones), and what that costs: ≈€45,500 per 100 MW of avoidable imbalance exposure at reBAP, concentrated in exactly the hours where failing is expensive.

Paper
KW-2026-01
Version
v1.1 analysis· history
Issued
2026-08-17
Author
Matthew Tanti
Version history
  • v0.1 · 2026-08-17 · working draft for review
  • v0.1.1 · 2026-09-01 · correction: an earlier version said Germany does not publish imbalance prices openly. It does. reBAP is published quarter-hourly by the four TSOs on netztransparenz.de. The Dutch leg carries the priced findings because it is the leg that was run, which is what the section now says.
  • v1.0 · 2026-09-02 · first public release. Content unchanged from 0.1.1; the version marks the paper as issued rather than circulating for review.
  • v1.1 · 2026-09-02 · the monetised leg is re-priced on German reBAP, published quarter-hourly by the four TSOs. The Dutch figures it previously carried are withdrawn: that series pair fails a forecast-integrity check (correlation 0.62, slope 0.61, forecast below actual at all 96 slots), so a euro figure scaled off mean load could not stand on it. Coverage results are unchanged.
Abstract

We audit every quarter-hour of German (DE-LU) day-ahead system forecasts from July 2025 through July 2026 (load, solar, onshore and offshore wind) against realized outcomes from the ENTSO-E Transparency Platform. Three results. First, the point forecasts carry conditional bias that vanishes in aggregate statistics: load is over-forecast by roughly 1.9-2.0 GW every midday. Second, the two standard ways of wrapping intervals around a point forecast, parametric Gaussian bands and rolling split-conformal quantiles, both undercover out-of-sample (83-89% at nominal 90%), and collapse regionally: solar's nominal-80% band covered 56-63% through spring 2026. Distribution shift breaks the exchangeability that conformal's guarantee rests on. Third, adaptive conformal inference restores marginal coverage at an honestly stated width cost, and reduces without removing the tendency to miss where a miss is expensive. Priced at German reBAP, trusting the Gaussian "90%" band carried roughly EUR 90,000 of un-reserved adverse exposure per 100 MW over the period, versus EUR 44,500 under the adaptive band, with misses concentrated in the intervals where imbalance was most expensive: the mean adverse spread was EUR 24.6/MWh overall and EUR 34.8/MWh on the intervals the band missed.

Background: pinned referencestranscluded at pinned version
  • Calibration background/calibration · v1.0 · tutorial

Introduction and notation the question · the data · notation

Every energy desk in Europe consumes the TSO day-ahead forecasts, and almost nobody asks the question this analysis answers: when those forecasts imply a range of outcomes, is the range honest? We audit thirteen months of German (DE-LU) day-ahead forecasts against what actually happened (every quarter-hour from July 2025 through July 2026) for system load, solar, onshore wind and offshore wind: 38,016 quarter-hours per series, worst-case missingness 0.05%, all from the ENTSO-E Transparency Platform.

The findings run in increasing order of consequence: a conditional bias that averages conceal (§ 02); the failure, out-of-sample, of both standard ways to build uncertainty bands, including the naive conformal one (§ 03); and the adaptive construction that keeps the average promise, at a stated price and without closing the gap where it is expensive (§ 04). § 05 records what this audit can and cannot show.

ea forecast error: actual minus forecast, in MW; negative = the forecast was too highCoveragethe fraction of the time an interval contained the outcome; a "90% band" claims coverage 90%αthe miss rate a method aims for; a 90% band targets α = 10%ACIadaptive conformal inference: a feedback rule that adjusts α after each miss (§ 04)
§ 01 / 05

The data, and the rules of evidence

Public forecasts, public outcomes, one strict discipline: every number scored strictly out-of-sample.

The Transparency Platform publishes each TSO’s day-ahead point forecasts alongside realized values. That pairing makes an audit possible without anyone’s permission: the forecast was committed before delivery, the outcome is measured after it, and the difference (the error e = actual − forecast) accumulates into a track record nobody curates.

Two rules govern everything below. First, any quantity fitted from history (a mean, a standard deviation, a quantile) uses a trailing 90-day window that ends the day before the observation it is scored on. No method ever sees its own exam. Second, every construction is evaluated per quarter-hour delivery slot, because grid errors have a schedule: a method that is right on average and wrong every noon should be caught, not excused.

§ 02 / 05

The bias clock

Germany's load forecast overshoots by two power stations' worth. Every day, at the same hours.

The aggregate numbers look respectable: load bias −385 MW on a ~55 GW system, solar −56 MW, offshore wind +176 MW. The clock tells another story.

Fig. 01: The bias clock Interactive · evidence

Mean forecast error by delivery hour (Europe/Berlin), 13 months of quarter-hours. Bars below the zero line mean the day-ahead forecast was too high at that hour. The overall mean (dashed) is what a summary statistic reports. The clock is what dispatch experiences.

-2300-1100+1100+2300mean -385 MWh0: +342 MWh1: +268 MWh2: +94 MWh3: +109 MWh4: +124 MWh5: +121 MWh6: +476 MWh7: +82 MWh8: -298 MWh9: -919 MWh10: -1526 MWh11: -1814 MWh12: -1967 MWh13: -1953 MWh14: -1804 MWh15: -1366 MWh16: -955 MWh17: -291 MWh18: +79 MWh19: +290 MWh20: +285 MWh21: +408 MWh22: +521 MWh23: +446 MW-1967-195306121823DELIVERY HOUR · MEAN ERROR, MW (ACTUAL − FORECAST)
Worst hour
h12: -1967 MW
Series mean bias
-385 MW
Ratio worst : mean
5.1×
ENTSO-E Transparency Platform · DE-LU · 2025-07-01 → 2026-07-31 · 38,016 quarter-hours/series · mean error by delivery hour

Load is over-forecast by 1.8–2.0 GW at hours 11–14: roughly two large power stations of demand that is predicted daily and daily fails to materialize, five times the size of the aggregate bias. The likely mechanism is well known (behind-the-meter solar eroding metered load faster than the forecast model learns); the size is something you only learn by measuring. Offshore wind leans the other way, under-forecast by ~240 MW through the night hours; solar’s midday under-forecast peaks around −420 MW.

Pitfall 2.1: The average forgivespit-2-1

A summary bias statistic averages over the clock. A desk that trades the midday blocks does not. Any audit, internal or vendor, that reports a single bias number per series has not yet said anything about the hours where the money is.

§ 03 / 05

Three ways to promise 90%

The parametric habit fails. So does the naive conformal alternative. Admitting that is the point.

The TSO publishes points, not ranges. Any desk wanting a range constructs it, and the standard construction is parametric: fit mean and standard deviation to recent errors, quote μ±zσ\mu \pm z\,\sigma. The textbook distribution-free alternative replaces the Gaussian fit with trailing empirical quantiles: a rolling split-conformal band. We ran both, identically windowed, strictly out-of-sample; and a third method introduced in § 04.

A desk with a quantitative team will say it does not build bands this way, and it is right. It runs quantile regression on weather features, or it buys a probabilistic forecast with bands already attached. That objection does not weaken the audit; it describes the normal case for it. What is under test is not how a band was built but whether the number printed on it is true. The procedure here scores whatever band you already quote, from whatever construction, against the outcomes you already have: it counts how often the band contained the answer, splits that count by regime, and prices the misses. Nothing in it needs to know the model. The two constructions audited below are in the paper because they are public and reproducible, not because they are the state of the art, and a better-built band is a more interesting subject for the same test rather than an exemption from it.

Fig. 02: Three ways to promise 90% Interactive · evidence

Coverage each method actually delivered at the selected promise, per series. All three use the same trailing 90-day windows per quarter-hour slot, scored strictly one day ahead. The madder segment is the gap between promise and delivery. Interval width is the honest price of closing it: shown in the readout.

52%57%63%69%74%80%85%91%96%promised 90%GaussianGaussian: 84.8% achieved84.8%rolling conformalrolling conformal: 83.5% achieved83.5%adaptive conformaladaptive conformal: 89.6% achieved89.6%ACHIEVED OUT-OF-SAMPLE COVERAGE
Gaussian width · mean / P95
7.5 / 11.1 GW
rolling conformal width · mean / P95
7.4 / 11.0 GW
adaptive conformal width · mean / P95
8.8 / 14.2 GW
Same panel · trailing 90-day windows per quarter-hour slot · scored one day ahead · madder = shortfall vs promise

At nominal 90%, the Gaussian band delivered 84.8–88.9% depending on series; rolling conformal delivered 83.2–86.6%. Neither is a 90% band. The aggregate conceals worse: solar’s nominal-80% construction covered 56–63% through spring 2026, precisely when new solar capacity made the error distribution grow faster than any trailing window could learn.

Pitfall 3.1: Conformal is not a magic wordpit-3-1

Split-conformal’s celebrated finite-sample guarantee is conditional on exchangeability: tomorrow’s error drawn from the same distribution as the window’s. A growing solar fleet is precisely a violation of that condition. Applied statically under distribution shift, conformal undercovers just like the Gaussian it replaced. A warranty is only as good as its conditions.

1This finding argues against our own field's naive pitch, which is why it leads. The value of conformal methods under shift is not the static guarantee; it is that the guarantee's failure is detectable and repairable, which § 04 demonstrates.
§ 04 / 05

The construction that keeps its promise

Adaptive conformal inference tracks its own misses and repairs them. Marginal coverage returns to nominal; the expensive hours stay harder.

Adaptive conformal inference (Gibbs & Candès, 2021) treats the miss rate as a control problem: after every observation, the target miss rate αt\alpha_t is nudged,

αt+1=αt+γ(αmisst),\alpha_{t+1} = \alpha_t + \gamma\,(\alpha - \mathrm{miss}_t),

so the very next interval widens after a miss and tightens after a long run of hits. The step size γ is a plain tuning constant (larger reacts faster and wobbles more), and it was fixed at γ = 0.02 before any evaluation, not tuned to the results. Same trailing windows, same no-lookahead discipline, same data.

The result, visible in Fig. 02 for every series and level: achieved coverage lands at 89.0–89.7% against a 90% promise, 79.3–79.9% against 80%, 93.1–94.6% against 95%. The price is stated rather than hidden, and stated in the tail, not just the mean, because an average width in an article about the sins of averaging would be an irony. The adaptive load band at 90% averages 8.8 GW against the Gaussian’s 7.5 GW; at the 95th percentile of widths it quotes 14.2 GW against the Gaussian’s 11.1 GW, and its single widest interval reached 19.3 GW (onshore wind’s worst: 25.8 GW). Those wide moments are the method telling you it recently missed and does not yet trust the window. That is information, not noise. The Gaussian’s narrower quotes at those same moments were never actually 90% bands. Calibration does not make uncertainty larger; it stops you understating it.

Where the aggregate table still flatters everyone is regime. The month-by-month cut is the figure this article’s own Recipe demands:

Fig. 03: Coverage through the regime break Interactive · evidence

The same three constructions, scored month by month. Aggregate tables average this away; the promise is only as good as its worst regime. The shaded band marks spring 2026, when German solar additions shifted the error distribution faster than a trailing window learns.

spring 202650%60%70%80%90%100%Gaussian2025-08 · Gaussian: 87.5% (n=192)2025-09 · Gaussian: 91.4% (n=2,880)2025-10 · Gaussian: 88.9% (n=2,972)2025-11 · Gaussian: 94.8% (n=2,880)2025-12 · Gaussian: 96.5% (n=2,976)2026-01 · Gaussian: 87.6% (n=2,976)2026-02 · Gaussian: 76.0% (n=2,688)2026-03 · Gaussian: 61.2% (n=2,972)2026-04 · Gaussian: 63.3% (n=2,880)2026-05 · Gaussian: 63.7% (n=2,976)2026-06 · Gaussian: 61.2% (n=2,880)2026-07 · Gaussian: 87.2% (n=2,976)rolling conformal2025-08 · rolling conformal: 82.3% (n=192)2025-09 · rolling conformal: 86.6% (n=2,880)2025-10 · rolling conformal: 85.5% (n=2,972)2025-11 · rolling conformal: 85.9% (n=2,880)2025-12 · rolling conformal: 82.2% (n=2,976)2026-01 · rolling conformal: 81.5% (n=2,976)2026-02 · rolling conformal: 71.3% (n=2,688)2026-03 · rolling conformal: 54.9% (n=2,972)2026-04 · rolling conformal: 55.6% (n=2,880)2026-05 · rolling conformal: 58.4% (n=2,976)2026-06 · rolling conformal: 61.1% (n=2,880)2026-07 · rolling conformal: 81.4% (n=2,976)adaptive conformal2025-08 · adaptive conformal: 81.8% (n=192)2025-09 · adaptive conformal: 86.8% (n=2,880)2025-10 · adaptive conformal: 82.5% (n=2,972)2025-11 · adaptive conformal: 82.1% (n=2,880)2025-12 · adaptive conformal: 79.2% (n=2,976)2026-01 · adaptive conformal: 77.1% (n=2,976)2026-02 · adaptive conformal: 74.0% (n=2,688)2026-03 · adaptive conformal: 65.2% (n=2,972)2026-04 · adaptive conformal: 74.4% (n=2,880)2026-05 · adaptive conformal: 79.1% (n=2,976)2026-06 · adaptive conformal: 82.0% (n=2,880)2026-07 · adaptive conformal: 89.0% (n=2,976)25/0825/1025/1226/0226/0426/06MONTH · OUT-OF-SAMPLE COVERAGE vs PROMISED 80%
Gaussian · worst month
61.2% (2026-06)
rolling conformal · worst month
54.9% (2026-03)
adaptive conformal · worst month
65.2% (2026-03)
Same panel and constructions · monthly out-of-sample coverage · shaded band = spring 2026 solar fleet shift

The month-by-month cut also keeps us honest about what adaptivity is. When the solar shift hits in March 2026, all three methods get hurt: ACI’s worst month is 65% against an 80% promise, because a reactive method cannot see a regime break coming. The difference is what happens next: ACI walks back to 74% in April, 79% in May, 82% by June, while the static constructions sit in the mid-50s to low-60s for four consecutive months, still quoting “80%” the whole time. Adaptivity is not immunity; it is recovery measured in weeks instead of quarters, and a stated miss you can see happening, instead of a silent one.

Recipe 4.1: Audit your own bandsrec-4-1
  1. Pair each committed forecast with its realized outcome; never re-fit after the fact.
  2. Score whatever band you quote (parametric or otherwise) strictly out-of-sample, per delivery slot.
  3. Report coverage by regime (hour, season, fleet state), not only in aggregate.
  4. If coverage drifts from the promise, adapt the miss rate, not the story.
§ 05 / 05

What this audit can and cannot show

TSO forecasts are the public benchmark, not your stack. The method transfers, the numbers will not.

What do these percentage points cost? German imbalance answers, because reBAP is published quarter-hourly by the four TSOs on netztransparenz.de. Pricing every escape from the trailing Gaussian “90%” band at that quarter-hour’s adverse imbalance-versus-day-ahead spread (favourable spreads clipped to zero: risk side only), scaled to a portfolio with 100 MW mean load whose errors move proportionally to the system’s (an assumption, stated): €90,000 of un-reserved adverse exposure over the 13 months, against €44,500 under the adaptive band: a ≈€45,500 difference per 100 MW. The stated band missed 15.2% of quarter-hours. And the property the averages hide: the mean adverse spread over all intervals was €24.6/MWh, while on the intervals the Gaussian band missed it was €34.8/MWh. The band fails disproportionately in exactly the intervals where failing is expensive. Miscalibration and market stress share causes; that correlation is the audit’s sharpest commercial fact.

Pitfall 4.1: an earlier version of this section priced the Dutch legpit-4-1

Until v1.1 this figure came from a Dutch run, because NL imbalance prices were pulled first. That leg has been withdrawn. The two ENTSO-E series behind it are not a forecast and its outcome: the published NL load forecast sits below realised load at every one of the 96 quarter-hour slots, correlating at 0.62 on a slope of 0.61, against 0.96 and 1.02 for the German pair. The coverage arithmetic survived that, because a trailing per-slot window absorbs a stable offset, but a euro figure scaled off mean load does not. The specimen audit now runs a series-integrity check before scoring anything, and keeps the Dutch run as the case where it fires.

Beyond that, this analysis can show whether a published point forecast is conditionally biased, and whether a band constructed around it holds its promised frequency. Those are checkable against outcomes. It cannot show whether your forecasts share these failure modes: portfolio forecasts differ from system forecasts, and a desk’s decision layer consumes uncertainty differently than a TSO’s.

One consequence deserves naming, because it is where this method has teeth. If you buy forecasts from a vendor who quotes confidence bands, everything above applies to them exactly as it applies to the TSO. Coverage is checkable by the customer, from data the customer already holds (committed forecasts and realized outcomes) without the vendor’s cooperation or their model. A vendor whose “95%” bands blow through weekly in a windy quarter is not suffering “unprecedented weather”; they are failing an audit nobody has run. Recipe 4.1 doubles as the four questions to put in your next RFP or quarterly vendor review.

The audit methodology (the pairing discipline, per-slot scoring, regime-conditional coverage, the adaptive repair) transfers unchanged to any stack that can export twelve months of forecasts and outcomes.

The engagement Forecast Calibration Audit

Forecast Calibration Audit: your committed forecasts and outcomes (or your vendor’s), audited exactly as above. Input: one CSV per series (timestamp, forecast, actual, stated band if any); 12+ months. You get: the bias clock and coverage ladder for your portfolio; coverage by regime and by season; the width-tail accounting; a monetized shortfall estimate against your market’s prices; an adaptive recalibration demonstration on your own data; a 45-minute readout. Terms: fixed price by portfolio size, two weeks from data delivery; runs on your infrastructure if preferred; your data is deleted on delivery and never used elsewhere; standard mutual NDA before anything moves. A scoped single-series Vendor Band Audit is available as a smaller first step; a specimen of that deliverable, run end to end on a public series, shows exactly what arrives. If your bands are honest, the audit proves that. A marketable result too.

Matthew Tanti · Kwantil, Malta · contact via kwantil.com

Cite this articlebibtex · plain text
@article{tanti2026,
  author  = {Matthew Tanti},
  title   = {German Grid Forecasts: Where the Uncertainty Bands Fail},
  journal = {Kwantil},
  year    = {2026},
  number  = {KW-2026-01},
  version = {1.1},
  url     = {https://kwantil.com/papers/german-grid-forecast-calibration/}
}

Matthew Tanti (2026). German Grid Forecasts: Where the Uncertainty Bands Fail. Kwantil, KW-2026-01, v1.1.

analysis · v1.1· Code and data ↗· Errors? Write in.