Writing

Defending a negative result

A positive result obtained under pressure can rest on one lucky split; a negative result attacked from three independent directions and left standing is the more robust piece of science. This is why Cascade's central finding, a measured boundary of predictability, is defended harder than the result that worked.

draft — being rewritten

There is a strong gradient pulling every result towards the positive. A model that clears its bar gets written up; a model that misses it gets quietly retuned until it clears, or dropped. The trouble is that a positive obtained that way is fragile. It can rest on one lucky split, one metric chosen after the outcome was known, one threshold nudged into place. A negative result that has been attacked from three independent directions and left standing is, perversely, the sturdier piece of science. Cascade’s central finding is one of those, and it is the part of the project I defend hardest.

The thing that worked, first

Cascade models how a weather shock propagates through the European energy system. The system is written as a typed graph (357 nodes / 675 edges), every edge carries the physical mechanism of its coupling in language, and a learned operator advances stress along the edges from six hours to fourteen days ahead. It is judged inside a curated catalogue of shock episodes (59 curated shock episodes), storms and heatwaves and supply crises and Dunkelflauten, rather than on the ordinary weather days that would flatter it exactly where it matters least.

There is a positive result, and it holds. On held-out crisis years the operator ranks stressed against quiet nodes well above persistence (0.738 vs 0.652 persistence), it clears persistence on the large majority of test episodes it never trained on, and removing the graph forfeits roughly half the margin. The propagation structure is doing real work, not decorating a per-node baseline.

Stress-ranking skill on held-out test years: the graph operator leads, its graph-free ablation and persistence trail, and climatology sits at the no-skill line. Stress-ranking skill on held-out test years: the graph operator leads, its graph-free ablation and persistence trail, and climatology sits at the no-skill line.
Stress-ranking skill on held-out test years. The operator beats persistence, and removing the graph forfeits about half its margin. Cascade

I lead with this so the negative that follows is not the alibi of a model that could not do anything. The operator works where the physics is autocorrelated. It is at the harder boundary that it stops.

The thing that did not

Onset of disruption, a node going from quiet to stressed for the first time, is not recoverable from node history on this graph (AUC 0.35 vs 0.65 pre-registered bar). The classifier lands well below a bar that was pre-registered and frozen before any test episode was scored, so the miss cannot be an after-the-fact goalpost moved into reach.

The December 2022 cold-still spell replayed on the Cascade graph, node colour encoding each node's daily peak stress against its own climatology as stress spreads across the continent.
The December 2022 cold-still spell, a held-out episode, replayed on the graph. The propagation is legible in hindsight; anticipating which quiet node lights up next is the part that fails. Cascade

A single failed metric is a weak claim, because a tuning miss looks identical to a real limit from the outside. So the onset question was put three independent ways: a quantile-median rule, an upper-tail exceedance rule, and a purpose-built class-weighted exceedance classifier built specifically to chase the rare positives. All three land below the bar. When three methods that share no machinery fail together, the failure is not a defect of any one of them. It is a property of the problem.

Why it is structural, not a tuning miss

The shared failure mode has a clean mechanism, and naming it is what turns “we could not do it” into “it cannot be done this way, and here is why”. Quietness predicts staying quiet. A node with no recent history of stress is, to any score conditioned on that history, the least likely to light up next, so every such score sorts the eventual spikers to the bottom of its ranking at precisely the moment they matter. The graph-neighbour signal that ought to override a node’s own autocorrelation is too faint on this system to do so.

The root is visible in the training objective. The operator is fitted with an anomaly-weighted pinball loss, which is variance-shrinking and pulls the quantiles inward. That same inward pull under-disperses the forecast (the central band covers well short of nominal, 66% vs 80% nominal) and scores a quiet node as safe, so the calibration failure and the onset failure are one mechanism seen twice. The boundary of predictability here is therefore not a data-volume problem to be trained away with a bigger model or a longer record. It is a consequence of conditioning on a node’s own quiet past, and it would survive any amount of the same kind of data.

Why the negative is the deliverable

Knowing where that boundary sits is the deliverable a risk carrier actually needs. Before anyone prices propagation risk, the first question is which parts of it are forecastable at all, and an honest “onset is not, on this graph, from this signal” answers it. An inflated headline that hides the limit does not make onset predictable; it just relocates the discovery to production, where a smelter or a balancing desk finds it instead. A measured negative closes the question. A fragile positive reopens it at the worst possible time.

There is a version of this project that reports only the stress-ranking win and files the onset study in a drawer. It would show better and be worth less. Defending a null takes more discipline than banking a positive, because the instinct is to read a null as a failure of the study rather than a finding about the world, and to keep tuning until the discomfort goes away. Triangulation is the discipline that resists it. Three independent attacks that all fail are not three disappointments. They are the evidence that the limit is real, and a real limit, precisely located, is a result.