Forecast Proof · TimesFM vs. PJM · Blanc Quant Services
Google's TimesFM reduced forecast error by 30% compared to the best-performing simple baseline model. However, it performed 37% worse than PJM's published day-ahead forecast. The adoption rules were frozen before either result was scored.
01 · Freeze the rules
Before the candidate produced a single forecast, we wrote and cryptographically hashed the full evaluation contract: dataset, horizon, baselines, folds, the primary metric, and the exact bar for adoption. Once frozen, the rules cannot move — a deliberate guard against the oldest failure in forecasting evaluation, which is deciding what counts as success after seeing the score.
The baselines were a ladder, not a strawman: last-value, 24-hour and 168-hour seasonal naive, exponential smoothing — and the incumbent: the day-ahead demand forecast the grid actually publishes in EIA-930. The incumbent comparison is mandatory in every Forecast Assurance evaluation.
02 · Test what already works
Every model was scored on identical timestamps. Shorter bar means lower error means better forecast.
PJM, TimesFM, and the simple baselines are drawn on one scale but did not share one information set — see the verdict below. The exact values are in the table.
| Forecaster | Mean seasonal MASE | Class |
|---|---|---|
| PJM published day-ahead forecast | 0.894 | incumbent · winner |
| TimesFM 2.5-200M, zero-shot | 1.226 | AI candidate |
| Seasonal naive (24h) | 1.741 | baseline |
| Seasonal naive (168h) | 2.502 | baseline |
| Last value | 2.743 | baseline |
| Exponential smoothing (span 24) | 2.755 | baseline |
03 · Issue the verdict
TimesFM demonstrated strong performance by outperforming a well-built seasonal naive model by 30% with zero training, using only on-demand historyin is the difference between guessing and planning.
But the two forecasts did not have equivalent information sets. PJM's operational forecast can draw on inputs such as weather that the tested zero-shot TimesFM configuration never received. So the result condemns this configuration, on this data and these timestamps, under this protocol — not the model in general, and PJM's forecast was not recreated internally. A TimesFM variant with weather covariates would be a different evaluation with a different frozen contract.
Forecast Assurance rejects configurations, not reputations.
A method that cannot deliver bad news is not assurance.
04 · Seal the evidence
Every number above traces to a sealed evidence package: the frozen contract, raw fold-level outputs, deterministic leakage and timestamp-alignment checks, and the computed metrics. Deterministic checks confirmed no future timestamps in any training window, identical evaluation timestamps for every model, no dropped observations, and that the incumbent is an exogenous forecast, not same-hour actuals.
An independent second model, running locally on the same workstation with zero data egress, reviewed the completed evaluation adversarially and ruled it valid, with the KILLED verdict standing at high confidence. The entire evaluation used about 2.5 seconds of model runtime on a single workstation — the protocol is the expensive part, never the computation. Our fee is never contingent on which model wins.
The evidence is organized as a sealed Forecast Assurance Evidence Package: the frozen contract, data manifest, candidate and incumbent forecasts, fold definitions, leakage and alignment checks, raw fold results, metric calculations, adversarial review, limitations, the machine-verifiable decision record, a hash manifest, and a reproduction runbook. The decision is regenerated by applying the frozen gate to the computed metrics — no analyst or stakeholder can promote a failing configuration by hand.
Objective forecast benchmarking is not new — public frameworks such as EPRI's Forecast Arbiter already support impartial, repeatable evaluation. What Blanc Quant Services adds is narrower: a preregistered, fixed-fee adoption decision against your own incumbent, sealed before any commercial pressure can rewrite it, delivered by an independent utility engineer.
Discuss a Forecast EvaluationRead the full Forecast Assurance overview, or email CEO@blancquantsystems.com.