Specimen · Vendor Band Audit

Forecast band calibration: DE-LU system load

A worked example of the deliverable, formatted exactly as a client receives it. It is run on a public series rather than on client data, and every number in it regenerates from the linked script and notebook.

Verdict

The two bands a desk builds without a calibration step do not deliver the coverage they state, and they fail hardest in the quarter-hours where being wrong is most expensive. Of 3 constructions audited, 2 carry Material findings and 1 does not fail on this evidence. The band that holds pays for it in width, and that price is scored rather than mentioned: at nominal 90% it quotes 8,820 MW on average against the parametric band's 7,547 MW, 17% wider on average, and never wide enough to be unusable on this series.

IDFindingSeverityExposure at 90%
F-01Parametric Gaussian band at nominal 90% covered 84.8%MaterialEUR 90,042
F-02Rolling split-conformal band at nominal 90% covered 83.5%MaterialEUR 94,400
F-03Adaptive conformal band at nominal 90% covered 89.6%ObservationEUR 44,495

Scope and exclusions

What an audit did not look at is the half a reader cannot infer, so it is stated rather than left to the method section.

What this audit did not check
  • Whether the forecast itself can be improved. This audits the stated uncertainty around it, not the point forecast.
  • Any series other than system load, and any period outside the 13 months above.
  • Whether the desk actually reserves flexibility equal to the stated band. The exposure figures assume it does; if it reserves more, they overstate.
  • Price risk beyond the imbalance-versus-day-ahead spread. No intraday trading, no portfolio effects, no cross-commodity exposure.
  • Whether the adaptive construction would hold on a future regime change. It is reactive by design and this report says so under F-03.

Findings register

F-01 Material

Parametric Gaussian band at nominal 90% covered 84.8%

Why this severity
nominal outside the coverage interval, and misses run 1.35x more frequent in the most expensive intervals (interval lower bound 1.16 > 1)
Measurement
Headline 90%: coverage 0.848, day-block bootstrap 95% interval [0.823, 0.871] over 32,252 scored intervals. Every audited level: 80%: covered 73.3% [70.3%, 76.2%], concentration 1.24x [1.10, 1.36], Significant; 90%: covered 84.8% [82.3%, 87.1%], concentration 1.35x [1.16, 1.55], Material; 95%: covered 91.2% [89.3%, 93.0%], concentration 1.50x [1.20, 1.81], Material
Consequence
EUR 90,042 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
Recommended action
Re-state the band at its measured coverage, or replace the parametric step. The failure is the Gaussian assumption itself: a symmetric interval at a fixed multiple of sigma (1.645 at 90%) cannot track an error distribution whose shape moves, and widening it uniformly buys coverage in the calm hours it already had.
Limit of this finding
Measured on this series over this period. The interval is block-bootstrapped by day and so carries within-day dependence; it does not carry uncertainty in the price series used to value the misses.

F-02 Material

Rolling split-conformal band at nominal 90% covered 83.5%

Why this severity
nominal outside the coverage interval, and misses run 1.29x more frequent in the most expensive intervals (interval lower bound 1.11 > 1)
Measurement
Headline 90%: coverage 0.835, day-block bootstrap 95% interval [0.810, 0.860] over 32,252 scored intervals. Every audited level: 80%: covered 72.7% [69.7%, 75.7%], concentration 1.25x [1.13, 1.39], Material; 90%: covered 83.5% [81.0%, 86.0%], concentration 1.29x [1.11, 1.50], Material; 95%: covered 89.3% [87.3%, 91.3%], concentration 1.43x [1.19, 1.71], Material
Consequence
EUR 94,400 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
Recommended action
Do not read the conformal guarantee as insurance here. It is conditional on exchangeability, and a trailing window under distribution shift violates that condition, which is why this construction covers no better than the parametric one it was meant to replace. Either recalibrate adaptively or widen on a measured schedule.
Limit of this finding
As above. Note also that this construction is the static split-conformal band, not the adaptive one: the finding says nothing about conformal methods that update, which are audited separately as F-03.

F-03 Observation

Adaptive conformal band at nominal 90% covered 89.6%

Why this severity
nominal coverage lies inside the day-block bootstrap interval; this evidence does not show the band failing at this level
Measurement
Headline 90%: coverage 0.896, day-block bootstrap 95% interval [0.877, 0.914] over 32,252 scored intervals. Every audited level: 80%: covered 79.4% [76.7%, 82.0%], concentration 1.31x [1.16, 1.48], Observation; 90%: covered 89.6% [87.7%, 91.4%], concentration 1.42x [1.19, 1.67], Observation; 95%: covered 94.3% [92.9%, 95.7%], concentration 1.53x [1.22, 1.91], Observation
Consequence
EUR 44,495 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
Recommended action
Nothing to remediate on coverage. Three things to watch instead. Its misses are the most price-concentrated of any construction audited, which the severity rule does not grade because concentration only enters once coverage has failed: the band holds its promise and still misses disproportionately when a miss is expensive, and the exposure column is where that shows. Beyond that: the band pays for its coverage in width, so read this row against the width line before reserving against it; and it is reactive, so a regime break costs a stretch of under-coverage before it recovers. If either matters, the question is how much width the desk can carry, not whether to recalibrate.
Limit of this finding
An Observation is not a pass in general. It says the stated level sits inside the bootstrap interval on this series and period, at a sample size this audit reports rather than hides. It does not establish that the band holds through a regime change, which this period contains only in part.

Evidence

Coverage against the level each band states

Intervals are 95% day-block bootstrap, resampling whole days. A band is only recorded as failing when the level it states falls outside that interval.

ConstructionStatesDelivered95% intervalVerdict
Parametric Gaussian80.0%73.3%[70.3%, 76.2%]fails
Parametric Gaussian90.0%84.8%[82.3%, 87.1%]fails
Parametric Gaussian95.0%91.2%[89.3%, 93.0%]fails
Rolling split-conformal80.0%72.7%[69.7%, 75.7%]fails
Rolling split-conformal90.0%83.5%[81.0%, 86.0%]fails
Rolling split-conformal95.0%89.3%[87.3%, 91.3%]fails
Adaptive conformal80.0%79.4%[76.7%, 82.0%]holds
Adaptive conformal90.0%89.6%[87.7%, 91.4%]holds
Adaptive conformal95.0%94.3%[92.9%, 95.7%]holds

Where the misses fall

An interval is expensive when its adverse imbalance spread lands in the top 20% of all scored intervals, which is EUR 39.97/MWh here. The threshold was fixed before this audit scored anything, though not in ignorance: the underlying study had already reported that misses cluster in costly intervals on this series in aggregate. What was not known in advance is any of the numbers below. The ratio below is the miss rate inside those intervals against the miss rate everywhere else.

ConstructionStatesMiss rate, expensiveMiss rate, restRatio95% interval
Parametric Gaussian80.0%31.5%25.5%1.24x[1.10, 1.36]
Parametric Gaussian90.0%19.2%14.2%1.35x[1.16, 1.55]
Parametric Gaussian95.0%12.0%8.0%1.50x[1.20, 1.81]
Rolling split-conformal80.0%32.5%26.0%1.2530x[1.13, 1.39]
Rolling split-conformal90.0%20.1%15.6%1.29x[1.11, 1.50]
Rolling split-conformal95.0%14.1%9.8%1.43x[1.19, 1.71]
Adaptive conformal80.0%25.3%19.4%1.31x[1.16, 1.48]
Adaptive conformal90.0%13.6%9.6%1.42x[1.19, 1.67]
Adaptive conformal95.0%7.9%5.1%1.53x[1.22, 1.91]

What the bands cost

For every quarter-hour the outcome escaped the stated band, the excess beyond the band edge is exposure the band said not to carry. It is priced at that interval's adverse imbalance spread and scaled to a 100 MW mean-load portfolio.

Construction80%90%95%
Parametric GaussianEUR 178,824EUR 90,042EUR 45,794
Rolling split-conformalEUR 183,883EUR 94,400EUR 55,411
Adaptive conformalEUR 124,307EUR 44,495EUR 21,339

A band that never misses does not exist, so the exposure figures need something to be read against. The reference below is an oracle: each slot's quantiles taken over the whole period, which no forecaster could have known. It is not a competitor and it is not in the register.

Band at nominal 90%CoverageMean widthMean missExposure
Oracle, with hindsight89.9%8,301 MW1,099 MWEUR 64,266
Adaptive conformal89.6%8,820 MW875 MWEUR 44,495

A band with hindsight coverage of 89.9% costs more than the adaptive band at 89.6%. That is not a broken baseline. Coverage counts misses; exposure pays for their size, and the two come apart for any band that changes width over time rather than only across slots. The oracle's width is fixed per slot for the whole period, so when a volatile stretch arrives it misses by a lot. The adaptive band widens after it misses, so it is wide when errors are large: mean miss 875 MW against 1,099 MW, on 6% more average width. No coverage statistic shows this.

Read the exposure column with that in mind, and with the width line below: exposure is non-increasing in width by construction, so any band can lower it by quoting a wider one.

The mean adverse spread across all scored intervals was EUR 24.6/MWh. Across the intervals the Gaussian 90% band missed it was EUR 34.8/MWh. The band fails disproportionately in the intervals where failing is expensive.

Series integrity, checked before anything is scored

An audit of a forecast band is worthless if the two series are not a forecast and its outcome. Every report opens with the same three checks, and this pair passes all three. The thresholds are fixed before a client's data arrives, but they were not chosen in ignorance: both columns below were already known when they were set, and they were placed to separate a pair like the control from a pair like this one. They are a tripwire, not a discovery.

CheckThis pairThresholdControl
Correlation, forecast against outcome0.958≥ 0.900.621
Slope of outcome on forecast1.0201 ± 0.150.610
Mean bias against mean load-0.71%± 5%15.14%
Verdictpassesfails

The control column is a real failure, not an illustration. It is the same audit run against Dutch system load, where the published forecast sits below realised load at every one of the 96 quarter-hour slots and the pair correlates at 0.62 with a slope of 0.61. No day-ahead forecast is wrong in one direction at every slot for thirteen months; a gap of that shape is definitional, and the two series are measuring different perimeters. We have not established which, and are not going to guess in a report.

That run is kept because a check nobody has seen fail is a check nobody should trust. It is also the answer to the obvious question about this section: it would have caught the problem on a client's data in the first hour, before anything downstream was built on it.

What a band costs in width

A band wider than 50% of mean load is not something an operator can reserve against, whatever its coverage. That fraction was fixed before scoring, like the expensive threshold, and the last column counts how often each construction crossed it. The adaptive band buys its coverage partly by going wide, and this is where that shows.

ConstructionStatesMean width95th pctWidestUnusable
Parametric Gaussian80.0%5,880 MW8,658 MW10,074 MW0.0%
Parametric Gaussian90.0%7,547 MW11,112 MW12,929 MW0.0%
Parametric Gaussian95.0%8,992 MW13,241 MW15,406 MW0.0%
Rolling split-conformal80.0%5,803 MW8,548 MW10,610 MW0.0%
Rolling split-conformal90.0%7,382 MW11,010 MW13,328 MW0.0%
Rolling split-conformal95.0%8,627 MW12,623 MW15,286 MW0.0%
Adaptive conformal80.0%6,804 MW10,971 MW18,443 MW0.0%
Adaptive conformal90.0%8,820 MW14,182 MW19,297 MW0.0%
Adaptive conformal95.0%10,530 MW16,323 MW19,567 MW0.0%

Assumptions register

Every number above rests on these. The direction column says which way the headline moves if the assumption is wrong, so a reader can discount it without having to guess.

AssumptionIf it is wrong
Portfolio errors move proportionally to system errors, scaled to 100 MW mean loadEither direction, and this is the single largest lever on the exposure figures. A portfolio less correlated with the system carries less exposure than stated; a concentrated one carries more.
Favourable spreads clipped to zero, risk side onlyInflates the exposure figures. A desk that earns on favourable imbalance would net some of this back.
The desk reserves flexibility equal to the stated bandInflates if the desk already reserves more than it states.
Trailing 90-day window, 60-day minimum, no lookaheadA longer window smooths regime breaks and would flatter the static constructions; a shorter one is noisier. Fixed before scoring.
Adaptive step size gamma = 0.02Fixed before evaluation, not tuned to this series. A larger step reacts faster and quotes wider.
Imbalance and day-ahead prices as published, unrevisedRevisions would move the exposure figures, not the coverage findings.
The forecast and outcome series describe the same perimeterChecked rather than assumed, and it holds here. This is the largest assumption a report like this carries silently, which is why it is a standing section above rather than a line in this table. The control column there shows what its failure looks like.
"Expensive" is defined over the audited period, after the factNeither direction, but it narrows the claim: the concentration result measures co-occurrence between misses and high spreads, not predictability. A desk cannot reserve against intervals it does not yet know are expensive, so this is evidence about where the risk sat, not a trading rule.

Method note

How severity is assigned

Severity is derived by rule, not asserted. The rule is fixed in the script and applied identically to every construction and level.

Material
The stated level falls outside the day-block bootstrap interval, and misses concentrate in expensive intervals by at least 1.25x, and that concentration interval excludes 1.
Significant
The stated level falls outside the interval, but the concentration test fails either limb: not distinguishable from none, or real but below the materiality floor.
Observation
The stated level falls inside the interval. The band holds at this level on this evidence, and sample size is allowed to win.

A construction takes the most severe rating it reaches at any audited level. The 1.25x floor is a commercial judgement rather than a statistical one: it stops a concentration of 1.02 with a tight interval being called Material. It is stated here so it can be argued with.

Where the floor bit

Parametric Gaussian at the 80.0% level concentrates 1.235x against a floor of 1.25x, and is recorded Significant rather than Material on that margin. A threshold chosen after seeing the results would not have been left where it costs the report a finding.

Why the intervals are bootstrapped and not binomial

Quarter-hour forecast errors are autocorrelated within a day and across neighbouring slots. Treating 32,252 intervals as independent trials would give a binomial interval far tighter than the evidence supports, and would promote findings that sample size cannot carry. Every interval here resamples whole days with replacement, 2,000 times, seed 20260901. That is why all three adaptive findings score Observation rather than a small win.

Lineage

DataENTSO-E Transparency Platform, zone DE: load forecast, load actual, day-ahead price, imbalance price
Scriptsrc/audit_specimen.py --zone DE
Codesrc/audit_specimen.py, sha256 a603d239f36cb5c212f04ece9904ec984c99d12679e1d821ce91084ca9c153a8
Artifact digestsha256 bb3068ff9cdac0930e3df1b37ef96a5534710bae5c7e3008f4646cfd051c1e61
Digest rulesha256 over every file in artifacts/specimen except meta.json, in filename order, feeding each file's name as UTF-8 followed by its bytes. meta.json is excluded because it carries the digest and cannot hash itself.
Outputsartifacts/specimen/: coverage.json, cost.json, findings.json, findings.csv, monthly.json, bias_slot.json, meta.json
Register as datafindings.csv carries all nine construction-level rows, including the levels this document summarises

A client engagement replaces the public series with your own and changes nothing else. This report is that script against public German data: system load from the ENTSO-E Transparency Platform, priced at reBAP as published by the four TSOs.