A gate that has never been run on data with no edge in it has an unknown false-positive rate. A kill verdict from a search with an unmeasured detection floor is a statement about the search, not about the market. Both were true of ours. This quarter we measured both. Neither answer flattered us.
On the false-positive side: run the whole pipeline — optimise, gate, forward — on synthetic series calibrated to a real market's moments and containing nothing to find. At full search width it returned a best-of-search survivor on every panel it was given, with manufactured terminal returns in the hundreds of per cent after costs.
On the false-negative side: plant a dislocation of known size and sweep it upward until the gate sees it. No cell in the first batch reached the conventional eighty-per-cent power standard at any dislocation we planted. The floor sits above what most contractual obligations pay.
And the disclosure that governs everything below. The generator producing these numbers has not yet passed its own out-of-sample moment-matching test. That table is owed as of this writing, and the campaign resting on it is under re-audit. So what follows is published as a measurement of our pipeline, not as a validated false-positive rate.
§ 1
Until this quarter, every control we ran was a real mechanic on real data: composition-matched, shuffled leg, swapped leader, sector-substituted. Such a control shares the market's genuine structure — trend, volatility clustering, cross-sectional dispersion, the lot. So a control that earns is not evidence of a broken treatment. It is evidence that the control was never a null. A momentum control in a trending quarter earns because momentum works in a trending quarter.
Formally, a matched control answers is this book better than other books like it? It cannot answer is this book better than nothing? We had been reading answers to the first question as though they answered the second. The distinction is set out in full, with the doctrine reversal it forced, in Two nulls, because they answer different questions (2026-08). Three incidents in one quarter, one cause:
The last of those has a second half worth publishing. Re-measured properly, those books were artefacts twice over: the pair was mis-specified against an unhedged benchmark, so roughly half the residual variance being "harvested" was uncontrolled currency basis, and one leg carried an unadjusted corporate action. The gate's rejection of them was correct — but for reasons nobody had established, on evidence nobody had checked. Being right by accident is not a working instrument.
So we built the market-level surrogate the literature asks for, and had never built. A regime-switching generator is fitted to a market's real moments and sampled to reproduce them — unconditional volatility, volatility-clustering persistence, fat tails, cross-asset correlation structure, the trading calendar — while containing no exploitable relationship. The surrogate-data requirement is explicit and is the standard we held it to: preserve the nuisance structure, destroy only the hypothesised signal. Twenty-four papers in our corpus bear on that construction.
Then run the pipeline on it. Across 6 independent market panels, the search returned a best-of-search "survivor" on every one. Terminal returns on the manufactured winners ran into the hundreds of per cent after costs, on data with nothing in it. The per-panel magnitudes are withheld: as an attributed table they map our coverage to our search's weakness.
A second pass rebuilt the null on each cell's own forward window rather than a common one, which is the harder test. Of the treatment arms carrying valid forward evidence, 0 of 10 cleared even the median of its own manufactured-survivor distribution. Two of the twelve cells carried no valid forward evidence at all — their winning parameterisation's lookback was longer than the window it was scored on — and their zero rows were placeholders, not measurements. We report that as a defect in the batch rather than as ten of twelve.
Reading
Printed beside the nineteen books that survived our specificity gate, the paragraph above invites an obvious inference: that the nineteen are manufactured. We would rather answer that here than leave it to a later page.
The rate is best-of-search on a single panel, measured before the specificity and forward legs run. The downstream nulls are what remove these, and they demonstrably remove some — a matched control ensemble caught a placebo beating its own original on two of three windows, which is exactly the job. What is being reported is not a broken gate.
Nor is it a clean bill. Turned on the surviving set itself, the same instrument left most of those books below their own manufactured-survivor bar: of the nineteen, 2 cleared, 3 were marginal and 14 sat below. Every approximation in that re-score made the bar easier to clear rather than harder, which is what makes the below-bar rows worth taking seriously and the clears worth replicating at full fidelity before anyone leans on them. We publish it as an open question rather than a verdict, because the instrument that produced it has not yet passed §3.
It is still the more expensive problem of the two, because it scales with search width and we control search width. That is §2.
§ 2
The obvious first hypothesis is that the search is simply too big. It is a good hypothesis and it is only half right. Cutting a panel's configuration count by an order of magnitude cut the manufactured-winner share by an order of magnitude too — and the bar it manufactures shrank without collapsing. Power against the original reference stayed near zero throughout. The failure is selection-limited, not search-size-limited. The two configuration counts are withheld; the ratio between them is the part that generalises.
The narrow instrument had to be repaired before it could be believed, which is itself the finding. On split-unadjusted closes the same narrow search manufactured a triple-digit-per-cent winner out of four corporate actions. Only after back-adjustment at read did its null bar mean anything. A related artefact in the same store: on unadjusted closes, two funds tracking the same index print an ex-dividend sawtooth into their spread, so the crossings a pair strategy trades are dividend-calendar events rather than dislocations. Those pairs outscored the genuine share-class mechanism they were benchmarked against.
So the response is not simply to search less. It is to name the mechanism before searching, and to price every trial you did try. Width is a decision about how many hypotheses you are willing to pay for, and every one of them raises the bar the real edge has to clear. That is the deflated-Sharpe argument stated operationally rather than as a footnote: the correction is not a post-hoc adjustment to a number, it is a constraint on how the number was allowed to be produced.
§ 3
One assumption carries this whole apparatus: that a calibrated generator can be edge-free in the way we need. If its regimes are mis-specified — too stationary, wrong correlation dynamics, a missing stylised fact — then a survivor on synthetic data proves nothing about real data, and a failure to survive proves nothing either. The generator becomes a new source of silent-zero, which is this house's dominant failure mode: a component that reports success while doing nothing.
The mitigation was written into the charter in the same paragraph as the instrument, because a mitigation named later is a mitigation not built. Fit the generator, sample it, and confirm out-of-sample that the synthetic series reproduces the real market's moments — volatility-clustering persistence, tail index, cross-sectional dispersion, autocorrelation at the horizons we actually trade. A generator that cannot pass its own moment-matching test is not a null, and any verdict resting on it is void.
We publish that rather than wait for it, because the alternative — printing a false-positive rate and quietly leaving its instrument unvalidated — is the failure this entire page exists to describe.
Two smaller instrument findings in the same spirit. Single-draw validation of the generator fails at window scale; the authoritative reference is an ensemble band across twenty-five draws, not one draw, and the single-draw form is retired. And in one re-scoring run a cell returned an explicit validation failure instead of a number: the books resting on it were scored on the remaining cells and the failure was carried loud through the record. An instrument that can say I don't know is worth more than one that always answers.
Which is why the synthetic null is registered in our survivability gate born-informational: it is computed and reported in every verdict, and its failure blocks nothing. Before it earned even that, it had to reproduce 3 of 3 documented verdicts on a book whose answer was already known. The pre-registered flip to a blocking leg waits on its proof against a full re-score of the incumbent book set. We do not let a new instrument set a floor on the strength of its own first results.
§ 4
The first generator buys a false-positive rate. A second one buys the number nobody measures: it generates the same calibrated series but injects a known relationship with tunable parameters — a cointegrating vector with specified half-life, band width and noise; a bounded process with a known post-lock drift; an episodically defended band of specified width and frequency — and sweeps the planted dislocation upward from zero until the gate starts finding it.
Batch one covered two mechanism classes across two market panels, sweeping residual scale against two mean-reversion half-lives with eight seeds per cell. No cell in the batch reached the conventional eighty-per-cent power standard anywhere in the sweep. Against the dislocation scales those two mechanisms actually offer, the search had little power at all. The power values, the minimum detectable edge and the detection floor in absolute units are withheld; the direction and the fact of measurement are published.
The operational consequence is a rule that now binds every verdict we issue, and it is retrospective rather than prospective-only:
A harness that judges everything must itself be judged, and by something other than its own output. The four checks that make it credible rather than decorative:
| Mechanism class | Null instrument | Power cell |
|---|---|---|
| Fund-rebalance mandate (geared and inverse tracking) | mean-reverting pair harness | measured, batch 1 |
| Defended band, episodic intervention | mean-reverting pair harness | measured, batch 1 |
| Event-driven classes (cohort events, cascade events, session-open dislocations) | own event-study nulls | routed out — the pair harness cannot see them by construction |
| All remaining classes | — | outstanding — a verdict citing one of these is a prerequisite batch, not a kill |
§ 5
The false-positive side of this gate is well instrumented. Its false-negative side was, until this quarter, unmeasured — and it is the expensive one. A false positive costs capacity. A false negative costs the edge, silently, and leaves a confident kill verdict in the record where an unanswered question should be.
A contractual obligation pays what it pays, and for most of them that is a small number. We enumerated 35 of them from primary documents — fund prospectuses and central-bank pages, each row carrying a verbatim-verified quote, 24 fund-rebalance mandates and 11 defended-band commitments — and measured 33 of them against our own cost model. 17 coverage holes are named in the same record, because a named absence is a finding and an unnamed one is a silent zero. With few exceptions the dislocation on offer sits below what a broad search can resolve. The dislocation sizes and the floor are both withheld in absolute units; printed together they would locate the floor as precisely as printing it.
The consequence for how search is allocated is uncomfortable enough to be worth stating plainly: a tournament is the moonshot channel, not the discovery channel. The corroboration is more uncomfortable still. Our best validated obligation-linked book was found by mechanism-first authoring — reading the obligation, pricing the dislocation, then writing the book — and not by the search that was supposed to find it. It sits below that search's floor and always did.
The floor exists and we have now measured it. Where it sits is proprietary, and is withheld: an absolute detection floor is a map of what we cannot see, which is worth more to a competitor than to a reader.
§ 6
That last item is the one we underestimated, and it deserves its own numbers. One whole-market daily equity feed in our store carries no delistings whatsoever across four years — not survivorship bias but survivorship absence, which independently blocks merger arbitrage, index deletions and every cross-sectional book on that feed. The corporate-action table behind it is empty, so its prices are unadjusted. The demonstrated cost of that one defect: a screen run against those unadjusted prices returned a positive mean net of costs across 111 trades, with the large majority of individual lines positive and test statistics clearing every conventional threshold — entirely a distribution sawtooth. 65% of its entries fell within five days of an ex-distribution drop against a 17% base rate. Re-run on a distribution-adjusted series, the same signal fires three times in four years and the mean turns negative. The two mean returns and the t-statistics are withheld as an attributable set — published together with the mechanism and the window they locate the screen; the counts, the base rates and the sign flip are the finding and are published.
Nothing about that screen was a research failure. It was a data defect wearing a t-statistic, and the pipeline had no instrument that could tell the difference until we built one.
Stands on
Cited by
Cross-references are hand-maintained; a link check runs before publish.
Revision history
Disclosure
This item publishes the direction, the counts, the apparatus and the policy. Withheld: the two configuration counts, the per-panel manufactured magnitudes, the power values, the minimum detectable edge, the detection floor and the census dislocations in absolute units, which market panels each cell was measured on, and the mean returns and t-statistics of the artefact screen in §6.