Specimen · Vendor Band Audit
Forecast band calibration: DE-LU system load
A worked example of the deliverable, formatted exactly as a client receives it. It is run on a public series rather than on client data, and every number in it regenerates from the linked script and notebook.
Verdict
The two bands a desk builds without a calibration step do not deliver the coverage they state, and they fail hardest in the quarter-hours where being wrong is most expensive. Of 3 constructions audited, 2 carry Material findings and 1 does not fail on this evidence. The band that holds pays for it in width, and that price is scored rather than mentioned: at nominal 90% it quotes 8,820 MW on average against the parametric band's 7,547 MW, 17% wider on average, and never wide enough to be unusable on this series.
| ID | Finding | Severity | Exposure at 90% |
|---|---|---|---|
| F-01 | Parametric Gaussian band at nominal 90% covered 84.8% | Material | EUR 90,042 |
| F-02 | Rolling split-conformal band at nominal 90% covered 83.5% | Material | EUR 94,400 |
| F-03 | Adaptive conformal band at nominal 90% covered 89.6% | Observation | EUR 44,495 |
Scope and exclusions
What an audit did not look at is the half a reader cannot infer, so it is stated rather than left to the method section.
- Whether the forecast itself can be improved. This audits the stated uncertainty around it, not the point forecast.
- Any series other than system load, and any period outside the 13 months above.
- Whether the desk actually reserves flexibility equal to the stated band. The exposure figures assume it does; if it reserves more, they overstate.
- Price risk beyond the imbalance-versus-day-ahead spread. No intraday trading, no portfolio effects, no cross-commodity exposure.
- Whether the adaptive construction would hold on a future regime change. It is reactive by design and this report says so under F-03.
Findings register
F-01 Material
Parametric Gaussian band at nominal 90% covered 84.8%
- Why this severity
- nominal outside the coverage interval, and misses run 1.35x more frequent in the most expensive intervals (interval lower bound 1.16 > 1)
- Measurement
- Headline 90%: coverage 0.848, day-block bootstrap 95% interval [0.823, 0.871] over 32,252 scored intervals. Every audited level: 80%: covered 73.3% [70.3%, 76.2%], concentration 1.24x [1.10, 1.36], Significant; 90%: covered 84.8% [82.3%, 87.1%], concentration 1.35x [1.16, 1.55], Material; 95%: covered 91.2% [89.3%, 93.0%], concentration 1.50x [1.20, 1.81], Material
- Consequence
- EUR 90,042 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
- Recommended action
- Re-state the band at its measured coverage, or replace the parametric step. The failure is the Gaussian assumption itself: a symmetric interval at a fixed multiple of sigma (1.645 at 90%) cannot track an error distribution whose shape moves, and widening it uniformly buys coverage in the calm hours it already had.
- Limit of this finding
- Measured on this series over this period. The interval is block-bootstrapped by day and so carries within-day dependence; it does not carry uncertainty in the price series used to value the misses.
F-02 Material
Rolling split-conformal band at nominal 90% covered 83.5%
- Why this severity
- nominal outside the coverage interval, and misses run 1.29x more frequent in the most expensive intervals (interval lower bound 1.11 > 1)
- Measurement
- Headline 90%: coverage 0.835, day-block bootstrap 95% interval [0.810, 0.860] over 32,252 scored intervals. Every audited level: 80%: covered 72.7% [69.7%, 75.7%], concentration 1.25x [1.13, 1.39], Material; 90%: covered 83.5% [81.0%, 86.0%], concentration 1.29x [1.11, 1.50], Material; 95%: covered 89.3% [87.3%, 91.3%], concentration 1.43x [1.19, 1.71], Material
- Consequence
- EUR 94,400 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
- Recommended action
- Do not read the conformal guarantee as insurance here. It is conditional on exchangeability, and a trailing window under distribution shift violates that condition, which is why this construction covers no better than the parametric one it was meant to replace. Either recalibrate adaptively or widen on a measured schedule.
- Limit of this finding
- As above. Note also that this construction is the static split-conformal band, not the adaptive one: the finding says nothing about conformal methods that update, which are audited separately as F-03.
F-03 Observation
Adaptive conformal band at nominal 90% covered 89.6%
- Why this severity
- nominal coverage lies inside the day-block bootstrap interval; this evidence does not show the band failing at this level
- Measurement
- Headline 90%: coverage 0.896, day-block bootstrap 95% interval [0.877, 0.914] over 32,252 scored intervals. Every audited level: 80%: covered 79.4% [76.7%, 82.0%], concentration 1.31x [1.16, 1.48], Observation; 90%: covered 89.6% [87.7%, 91.4%], concentration 1.42x [1.19, 1.67], Observation; 95%: covered 94.3% [92.9%, 95.7%], concentration 1.53x [1.22, 1.91], Observation
- Consequence
- EUR 44,495 of un-reserved adverse exposure per 100 MW mean load over the audited period at the 90% level
- Recommended action
- Nothing to remediate on coverage. Three things to watch instead. Its misses are the most price-concentrated of any construction audited, which the severity rule does not grade because concentration only enters once coverage has failed: the band holds its promise and still misses disproportionately when a miss is expensive, and the exposure column is where that shows. Beyond that: the band pays for its coverage in width, so read this row against the width line before reserving against it; and it is reactive, so a regime break costs a stretch of under-coverage before it recovers. If either matters, the question is how much width the desk can carry, not whether to recalibrate.
- Limit of this finding
- An Observation is not a pass in general. It says the stated level sits inside the bootstrap interval on this series and period, at a sample size this audit reports rather than hides. It does not establish that the band holds through a regime change, which this period contains only in part.
Evidence
Coverage against the level each band states
Intervals are 95% day-block bootstrap, resampling whole days. A band is only recorded as failing when the level it states falls outside that interval.
| Construction | States | Delivered | 95% interval | Verdict |
|---|---|---|---|---|
| Parametric Gaussian | 80.0% | 73.3% | [70.3%, 76.2%] | fails |
| Parametric Gaussian | 90.0% | 84.8% | [82.3%, 87.1%] | fails |
| Parametric Gaussian | 95.0% | 91.2% | [89.3%, 93.0%] | fails |
| Rolling split-conformal | 80.0% | 72.7% | [69.7%, 75.7%] | fails |
| Rolling split-conformal | 90.0% | 83.5% | [81.0%, 86.0%] | fails |
| Rolling split-conformal | 95.0% | 89.3% | [87.3%, 91.3%] | fails |
| Adaptive conformal | 80.0% | 79.4% | [76.7%, 82.0%] | holds |
| Adaptive conformal | 90.0% | 89.6% | [87.7%, 91.4%] | holds |
| Adaptive conformal | 95.0% | 94.3% | [92.9%, 95.7%] | holds |
Where the misses fall
An interval is expensive when its adverse imbalance spread lands in the top 20% of all scored intervals, which is EUR 39.97/MWh here. The threshold was fixed before this audit scored anything, though not in ignorance: the underlying study had already reported that misses cluster in costly intervals on this series in aggregate. What was not known in advance is any of the numbers below. The ratio below is the miss rate inside those intervals against the miss rate everywhere else.
| Construction | States | Miss rate, expensive | Miss rate, rest | Ratio | 95% interval |
|---|---|---|---|---|---|
| Parametric Gaussian | 80.0% | 31.5% | 25.5% | 1.24x | [1.10, 1.36] |
| Parametric Gaussian | 90.0% | 19.2% | 14.2% | 1.35x | [1.16, 1.55] |
| Parametric Gaussian | 95.0% | 12.0% | 8.0% | 1.50x | [1.20, 1.81] |
| Rolling split-conformal | 80.0% | 32.5% | 26.0% | 1.2530x | [1.13, 1.39] |
| Rolling split-conformal | 90.0% | 20.1% | 15.6% | 1.29x | [1.11, 1.50] |
| Rolling split-conformal | 95.0% | 14.1% | 9.8% | 1.43x | [1.19, 1.71] |
| Adaptive conformal | 80.0% | 25.3% | 19.4% | 1.31x | [1.16, 1.48] |
| Adaptive conformal | 90.0% | 13.6% | 9.6% | 1.42x | [1.19, 1.67] |
| Adaptive conformal | 95.0% | 7.9% | 5.1% | 1.53x | [1.22, 1.91] |
What the bands cost
For every quarter-hour the outcome escaped the stated band, the excess beyond the band edge is exposure the band said not to carry. It is priced at that interval's adverse imbalance spread and scaled to a 100 MW mean-load portfolio.
| Construction | 80% | 90% | 95% |
|---|---|---|---|
| Parametric Gaussian | EUR 178,824 | EUR 90,042 | EUR 45,794 |
| Rolling split-conformal | EUR 183,883 | EUR 94,400 | EUR 55,411 |
| Adaptive conformal | EUR 124,307 | EUR 44,495 | EUR 21,339 |
A band that never misses does not exist, so the exposure figures need something to be read against. The reference below is an oracle: each slot's quantiles taken over the whole period, which no forecaster could have known. It is not a competitor and it is not in the register.
| Band at nominal 90% | Coverage | Mean width | Mean miss | Exposure |
|---|---|---|---|---|
| Oracle, with hindsight | 89.9% | 8,301 MW | 1,099 MW | EUR 64,266 |
| Adaptive conformal | 89.6% | 8,820 MW | 875 MW | EUR 44,495 |
A band with hindsight coverage of 89.9% costs more than the adaptive band at 89.6%. That is not a broken baseline. Coverage counts misses; exposure pays for their size, and the two come apart for any band that changes width over time rather than only across slots. The oracle's width is fixed per slot for the whole period, so when a volatile stretch arrives it misses by a lot. The adaptive band widens after it misses, so it is wide when errors are large: mean miss 875 MW against 1,099 MW, on 6% more average width. No coverage statistic shows this.
Read the exposure column with that in mind, and with the width line below: exposure is non-increasing in width by construction, so any band can lower it by quoting a wider one.
The mean adverse spread across all scored intervals was EUR 24.6/MWh. Across the intervals the Gaussian 90% band missed it was EUR 34.8/MWh. The band fails disproportionately in the intervals where failing is expensive.
Series integrity, checked before anything is scored
An audit of a forecast band is worthless if the two series are not a forecast and its outcome. Every report opens with the same three checks, and this pair passes all three. The thresholds are fixed before a client's data arrives, but they were not chosen in ignorance: both columns below were already known when they were set, and they were placed to separate a pair like the control from a pair like this one. They are a tripwire, not a discovery.
| Check | This pair | Threshold | Control |
|---|---|---|---|
| Correlation, forecast against outcome | 0.958 | ≥ 0.90 | 0.621 |
| Slope of outcome on forecast | 1.020 | 1 ± 0.15 | 0.610 |
| Mean bias against mean load | -0.71% | ± 5% | 15.14% |
| Verdict | passes | fails |
The control column is a real failure, not an illustration. It is the same audit run against Dutch system load, where the published forecast sits below realised load at every one of the 96 quarter-hour slots and the pair correlates at 0.62 with a slope of 0.61. No day-ahead forecast is wrong in one direction at every slot for thirteen months; a gap of that shape is definitional, and the two series are measuring different perimeters. We have not established which, and are not going to guess in a report.
That run is kept because a check nobody has seen fail is a check nobody should trust. It is also the answer to the obvious question about this section: it would have caught the problem on a client's data in the first hour, before anything downstream was built on it.
What a band costs in width
A band wider than 50% of mean load is not something an operator can reserve against, whatever its coverage. That fraction was fixed before scoring, like the expensive threshold, and the last column counts how often each construction crossed it. The adaptive band buys its coverage partly by going wide, and this is where that shows.
| Construction | States | Mean width | 95th pct | Widest | Unusable |
|---|---|---|---|---|---|
| Parametric Gaussian | 80.0% | 5,880 MW | 8,658 MW | 10,074 MW | 0.0% |
| Parametric Gaussian | 90.0% | 7,547 MW | 11,112 MW | 12,929 MW | 0.0% |
| Parametric Gaussian | 95.0% | 8,992 MW | 13,241 MW | 15,406 MW | 0.0% |
| Rolling split-conformal | 80.0% | 5,803 MW | 8,548 MW | 10,610 MW | 0.0% |
| Rolling split-conformal | 90.0% | 7,382 MW | 11,010 MW | 13,328 MW | 0.0% |
| Rolling split-conformal | 95.0% | 8,627 MW | 12,623 MW | 15,286 MW | 0.0% |
| Adaptive conformal | 80.0% | 6,804 MW | 10,971 MW | 18,443 MW | 0.0% |
| Adaptive conformal | 90.0% | 8,820 MW | 14,182 MW | 19,297 MW | 0.0% |
| Adaptive conformal | 95.0% | 10,530 MW | 16,323 MW | 19,567 MW | 0.0% |
Assumptions register
Every number above rests on these. The direction column says which way the headline moves if the assumption is wrong, so a reader can discount it without having to guess.
| Assumption | If it is wrong |
|---|---|
| Portfolio errors move proportionally to system errors, scaled to 100 MW mean load | Either direction, and this is the single largest lever on the exposure figures. A portfolio less correlated with the system carries less exposure than stated; a concentrated one carries more. |
| Favourable spreads clipped to zero, risk side only | Inflates the exposure figures. A desk that earns on favourable imbalance would net some of this back. |
| The desk reserves flexibility equal to the stated band | Inflates if the desk already reserves more than it states. |
| Trailing 90-day window, 60-day minimum, no lookahead | A longer window smooths regime breaks and would flatter the static constructions; a shorter one is noisier. Fixed before scoring. |
| Adaptive step size gamma = 0.02 | Fixed before evaluation, not tuned to this series. A larger step reacts faster and quotes wider. |
| Imbalance and day-ahead prices as published, unrevised | Revisions would move the exposure figures, not the coverage findings. |
| The forecast and outcome series describe the same perimeter | Checked rather than assumed, and it holds here. This is the largest assumption a report like this carries silently, which is why it is a standing section above rather than a line in this table. The control column there shows what its failure looks like. |
| "Expensive" is defined over the audited period, after the fact | Neither direction, but it narrows the claim: the concentration result measures co-occurrence between misses and high spreads, not predictability. A desk cannot reserve against intervals it does not yet know are expensive, so this is evidence about where the risk sat, not a trading rule. |
Method note
How severity is assigned
Severity is derived by rule, not asserted. The rule is fixed in the script and applied identically to every construction and level.
- Material
- The stated level falls outside the day-block bootstrap interval, and misses concentrate in expensive intervals by at least 1.25x, and that concentration interval excludes 1.
- Significant
- The stated level falls outside the interval, but the concentration test fails either limb: not distinguishable from none, or real but below the materiality floor.
- Observation
- The stated level falls inside the interval. The band holds at this level on this evidence, and sample size is allowed to win.
A construction takes the most severe rating it reaches at any audited level. The 1.25x floor is a commercial judgement rather than a statistical one: it stops a concentration of 1.02 with a tight interval being called Material. It is stated here so it can be argued with.
Parametric Gaussian at the 80.0% level concentrates 1.235x against a floor of 1.25x, and is recorded Significant rather than Material on that margin. A threshold chosen after seeing the results would not have been left where it costs the report a finding.
Why the intervals are bootstrapped and not binomial
Quarter-hour forecast errors are autocorrelated within a day and across neighbouring slots. Treating 32,252 intervals as independent trials would give a binomial interval far tighter than the evidence supports, and would promote findings that sample size cannot carry. Every interval here resamples whole days with replacement, 2,000 times, seed 20260901. That is why all three adaptive findings score Observation rather than a small win.
Lineage
| Data | ENTSO-E Transparency Platform, zone DE: load forecast, load actual, day-ahead price, imbalance price |
|---|---|
| Script | src/audit_specimen.py --zone DE |
| Code | src/audit_specimen.py, sha256 a603d239f36cb5c212f04ece9904ec984c99d12679e1d821ce91084ca9c153a8 |
| Artifact digest | sha256 bb3068ff9cdac0930e3df1b37ef96a5534710bae5c7e3008f4646cfd051c1e61 |
| Digest rule | sha256 over every file in artifacts/specimen except meta.json, in filename order, feeding each file's name as UTF-8 followed by its bytes. meta.json is excluded because it carries the digest and cannot hash itself. |
| Outputs | artifacts/specimen/: coverage.json, cost.json, findings.json, findings.csv, monthly.json, bias_slot.json, meta.json |
| Register as data | findings.csv carries all nine construction-level rows, including the levels this document summarises |
A client engagement replaces the public series with your own and changes nothing else. This report is that script against public German data: system load from the ENTSO-E Transparency Platform, priced at reBAP as published by the four TSOs.