AUGUR_

by ELOQUENTIX · METHODOLOGY NOTES

How to benchmark an energy forecast honestly

Eloquentix Inc. · updated July 2026 · the working notes behind augur.eloquentix.com

Every forecasting vendor claims to be accurate. Almost none of the claims are comparable, and many are meaningless — not because the numbers are fake, but because the baseline was chosen to flatter. These are the rules we hold our own forecasts to. They are not original; they are just rarely all applied at once.

1. The baseline decides whether your number means anything

"Our model beats persistence by 72%" sounds impressive. For weather-driven generation it is close to content-free: persistence (tomorrow equals today) is such a weak day-ahead predictor of wind that any gradient-boosted model with a weather feed beats it by 70%+. We know because we measured exactly that on German wind — and then refused to quote it.

The only baseline that carries information is the strongest forecast the buyer already has for free. In European power markets that is the transmission system operator's own published forecast (in Germany via SMARD, in Belgium via Elia Open Data). Our headline numbers are always: our RMSE vs the TSO's RMSE, identical delivery slots, same information cutoff. Sometimes we lose — German day-ahead solar, where the TSO is genuinely excellent, reads −5.8% on our own dashboard. Publishing the losses is what makes the wins worth reading.

2. No lookahead, or the backtest is fiction

The cardinal sin of forecasting backtests is using information that did not exist at prediction time. The two common leaks:

A subtler variant bit us in production: a conformal calibration window that overlapped the training data. In-sample residuals understated the interval widths and live coverage collapsed to 26–46% against an 80% target. The fix — a calibration window fully disjoint from the fit — widened one model's bands by a full gigawatt. If your intervals were calibrated on data the model trained on, they are not calibrated.

The vintage leak (July 2026 — our most expensive lesson). Forecast archives keyed by valid time can stitch together the freshest run available for each hour — a run issued after your forecast origin at long horizons. Using such an archive is not reanalysis, and it passes every "forecast, not realized" check — yet it quietly hands the model a peek at near-delivery weather. Our internal audit proved this inflated our long-horizon solar results: rebuilt with run-vintage control (only forecast runs issued before the origin), the apparent day-ahead solar edge disappeared, and the wind edge halved but survived. Every number we now publish is vintage-controlled, and the corrected results replaced the old claims on this site. If your backtest uses a stitched forecast archive and does not pin the run initialization time, your long-horizon skill is probably fiction too.

3. Metrics that survive contact with other people's numbers

MAPE is the wrong metric for solar. It divides by the actual value, which is near zero at dawn and dusk, so the metric explodes exactly where forecasts are easiest in absolute terms. We report RMSE normalized by installed capacity (nRMSE/cap), which is stable and comparable to the published literature — good day-ahead wind models sit around 5–8% nRMSE/cap. Ours reads 6.1% on 50Hertz-area wind. When someone insists on MAPE we quote it on slots above 10% of capacity, labeled as such.

4. Probabilistic honesty: coverage is measurable

A point forecast answers "what's your best guess." Real decisions — how much reserve to hold, how much to bid — need "how wrong might you be." We publish P10/P50/P90 bands and measure empirical coverage: an 80% band should contain the realized value 80% of the time, on held-out data. Two things that turned out to matter in practice:

5. Claim skill only where you have it

Our intraday models ingest live grid measurements the TSO's products haven't absorbed yet. That information advantage is real but short-lived: we beat the TSO's continuously updated forecast by 40–78% at 15 minutes ahead, decaying to parity around three hours. Beyond that, the TSO's weather-model-driven forecast wins — so beyond that, our system defers to a calibrated blend of the operator's own products instead of pretending. The per-horizon crossover was measured, not assumed; we once shipped a 4-hour model that quietly performed worse than the forecast it replaced, and found it only by checking every horizon against the incumbent separately.

6. Live beats backtested

Backtests, however clean, are still the researcher grading their own exam. Since June 2026 Augur also runs live: an unattended system submits hourly probabilistic forecasts (P10/P50/P90 at 15-minute resolution) for Belgian solar and 50Hertz-area wind and solar to Predico, the collaborative forecasting platform built by INESC TEC and used by Elia, Belgium's grid operator — every gate, around the clock — and self-scores each submission against the operator's published forecast once actuals settle. Live operation is where honest methodology stops being philosophy: missed deadlines, stale data feeds and platform outages all score against you, exactly as they would against a production trading desk.

As of July 2026, Augur competes on Predico's production instance — promoted from the evaluation environment by the platform operator on the basis of measured forecast performance, and eligible for the monthly rewards paid to the best-performing forecasters. Rankings there are computed by the market operator, not by us — which is precisely the point.

The checklist