Forecast Proof · TimesFM vs. PJM · Blanc Quant Services

The AI forecast looked impressive. We still rejected it.

Google's TimesFM reduced forecast error by 30% compared to the best-performing simple baseline model. However, it performed 37% worse than PJM's published day-ahead forecast. The adoption rules were frozen before either result was scored.

01 · Freeze the rules

The rules came first

Before the candidate produced a single forecast, we wrote and cryptographically hashed the full evaluation contract: dataset, horizon, baselines, folds, the primary metric, and the exact bar for adoption. Once frozen, the rules cannot move — a deliberate guard against the oldest failure in forecasting evaluation, which is deciding what counts as success after seeing the score.

CandidateTimesFM 2.5-200M, zero-shot, univariate
Data2,000 contiguous hours, PJM demand, public EIA-930
Horizon24 hours
Folds5 rolling-origin, non-overlapping
MetricSeasonal MASE (m=24)
Adoption gateBeat strongest baseline ≥5%, win ≥4 of 5 folds, no fold >20% worse
contract sha2562aae637c257e93d4ea15756880d510803e10f25a2b4def78f8c91bb5bd76f651
data snapshot sha2569c11018e062c326ef6a0f009af6dbd20a41d63f52277dd8773a4ef1f28063b06

The baselines were a ladder, not a strawman: last-value, 24-hour and 168-hour seasonal naive, exponential smoothing — and the incumbent: the day-ahead demand forecast the grid actually publishes in EIA-930. The incumbent comparison is mandatory in every Forecast Assurance evaluation.

02 · Test what already works

Lower error is unmistakable

Every model was scored on identical timestamps. Shorter bar means lower error means better forecast.

PJM forecast 0.894 ★ winner TimesFM (AI) 1.226 seasonal naive 24h 1.741 seasonal naive 168h 2.502 last value 2.743 exp. smoothing 2.755 0.0 mean seasonal MASE → higher = worse

PJM, TimesFM, and the simple baselines are drawn on one scale but did not share one information set — see the verdict below. The exact values are in the table.

Mean seasonal MASE across five rolling-origin folds. Lower is better.
Result table (available even if the chart does not render)
ForecasterMean seasonal MASEClass
PJM published day-ahead forecast0.894incumbent · winner
TimesFM 2.5-200M, zero-shot1.226AI candidate
Seasonal naive (24h)1.741baseline
Seasonal naive (168h)2.502baseline
Last value2.743baseline
Exponential smoothing (span 24)2.755baseline

03 · Issue the verdict

The verdict

Rejected The tested configuration did not meet the preregistered adoption gate. TimesFM finished roughly 30% ahead of the best simple baseline — genuinely strong — and roughly 37% behind the incumbent, winning only 2 of 5 folds where 4 were required. This configuration did not justify replacing the incumbent.

TimesFM demonstrated strong performance by outperforming a well-built seasonal naive model by 30% with zero training, using only on-demand historyin is the difference between guessing and planning.

But the two forecasts did not have equivalent information sets. PJM's operational forecast can draw on inputs such as weather that the tested zero-shot TimesFM configuration never received. So the result condemns this configuration, on this data and these timestamps, under this protocol — not the model in general, and PJM's forecast was not recreated internally. A TimesFM variant with weather covariates would be a different evaluation with a different frozen contract.

Forecast Assurance rejects configurations, not reputations.

A method that cannot deliver bad news is not assurance.

04 · Seal the evidence

Reproducible, reviewed, sealed

Every number above traces to a sealed evidence package: the frozen contract, raw fold-level outputs, deterministic leakage and timestamp-alignment checks, and the computed metrics. Deterministic checks confirmed no future timestamps in any training window, identical evaluation timestamps for every model, no dropped observations, and that the incumbent is an exogenous forecast, not same-hour actuals.

An independent second model, running locally on the same workstation with zero data egress, reviewed the completed evaluation adversarially and ruled it valid, with the KILLED verdict standing at high confidence. The entire evaluation used about 2.5 seconds of model runtime on a single workstation — the protocol is the expensive part, never the computation. Our fee is never contingent on which model wins.

The evidence is organized as a sealed Forecast Assurance Evidence Package: the frozen contract, data manifest, candidate and incumbent forecasts, fold definitions, leakage and alignment checks, raw fold results, metric calculations, adversarial review, limitations, the machine-verifiable decision record, a hash manifest, and a reproduction runbook. The decision is regenerated by applying the frozen gate to the computed metrics — no analyst or stakeholder can promote a failing configuration by hand.

contract sha2562aae637c257e93d4ea15756880d510803e10f25a2b4def78f8c91bb5bd76f651
data snapshot sha2569c11018e062c326ef6a0f009af6dbd20a41d63f52277dd8773a4ef1f28063b06
results sha25681b346ddb5acfdc17ca69cd21390e6514fce8a7b040b3d48c03e4f9c165f3c9d
decision recordverdict: REJECT · frozen-gate parity confirmed · sealed evidence taxonomy: KILLED

Test your forecast before you adopt it

Objective forecast benchmarking is not new — public frameworks such as EPRI's Forecast Arbiter already support impartial, repeatable evaluation. What Blanc Quant Services adds is narrower: a preregistered, fixed-fee adoption decision against your own incumbent, sealed before any commercial pressure can rewrite it, delivered by an independent utility engineer.

Discuss a Forecast Evaluation

Read the full Forecast Assurance overview, or email CEO@blancquantsystems.com.