A matched control asks whether a book beats its neighbours. Only an edge-free generator asks whether it beats nothing at all. We had been reading answers to the first question as though they answered the second, and three separate incidents in one campaign were the bill for it.
This paper sets out the distinction, the three incidents that forced it, the instrument built in response, and the condition we impose on that instrument. It also records a reversal. In July we wrote that synthetic price paths stay out of edge decisions permanently. Inside the same month a calibrated synthetic null was written into the gate. Both positions are correct, because they answer different questions.
One thing should be said before anything else here is read as settled: the generator's own out-of-sample moment-match table is outstanding as of this writing, and the campaign resting on it is under re-audit by people who did not build it. A generator that has not passed its own test is not yet a null. We publish the instrument and its unfinished business together.
§1
Every control we run is a real mechanic on real data. Composition-matched baskets, a shuffled leg, a swapped leader, a substituted sector. That construction is deliberate, and it is what makes the control fair: it holds everything constant except the mechanism the candidate claims.
It also means the control shares the market's genuine structure — trend, volatility clustering, cross-sectional dispersion, the lot. A momentum control in a trending quarter earns, because momentum works in a trending quarter. So a control that earns is not evidence of a broken treatment. It is evidence that the control was never a null.
Stated formally, the composition null answers is this book better than other books like it? It cannot answer is this book better than nothing? The first is a relative benchmark. For years we had only the relative one, and we read it as though it were absolute.
This is not an argument against matched controls. Inside a book's own symbol set the surrogate is built correctly and it bites. A permutation null that keeps a winner's exact exposure profile — instruments, directions, durations, sizes — and randomises only when each trade was placed, 500 times per book, found that 25 of 182 measurable graded winners could not beat random timing of their own trades at p ≥ 0.3, against 93 that cleared at p ≤ 0.05. The technique is not the failure. The failure is that we never built the market-level surrogate.
§2
All three landed inside one campaign, on three unrelated book families, and each was read at the time as a local defect. They are one defect.
In one forward panel, control- and placebo-named books took 7 of the 10 highest outcomes. They were 35% of the books in that panel, so the base-rate expectation was 3.5 of 10. Their median forward outcome beat the treatment books' median in 3 of 3 panels tested (Fig. 1). Under the reading we had been using, that is a fleet-wide indictment of the treatments. Under the correct reading it is a statement about the controls: they were never null, so they were never a floor.
A sector placebo — built to destroy the shared mechanism its treatment claimed — passed on evidence indistinguishable from the treatment's, and a later re-score put the pair at the same edge. The composition-matched control could not separate the claimed mechanism from ordinary sector beta, because the control carried the same exposure and therefore inherited the same result. The treatment was retired. The p-value attaches to a live-adjacent book set and is withheld; the direction and the verdict are published.
A set of books entered as nulls turned out to carry prospectus-mandated daily rebalances. A named obligated party, a dated trigger, a contractual obligation to transact: that is our own definition of an enforced mechanism, and a prospectus rebalance mandate is listed on our own coverage page as a canonical example of one. They were not nulls. They were untested treatments sitting in the control arm, and the registry field that should have auto-tagged them as mechanism-free never fired — a silent success return, which is the failure mode this house engineers hardest against.
A claim had been resting on those books: that a strong risk-adjusted result is the default output of optimisation rather than a discovery. It had been repeated into our own memory, our own analysis and our own conversation. It rested on those books being nulls. They were not, so the claim was withdrawn the day they were re-read. It may still be true. It has never been tested, and we no longer say it.
Re-measured properly, the arm's apparent result was a data artefact twice over. Roughly half the residual variance the books were "harvesting" was an uncontrolled currency-hedge basis rather than tracking error, and an unadjusted corporate action sat in the store underneath the series. The gate's original rejection of them was correct — but for reasons nobody had established, on evidence nobody had checked. A correct verdict reached by a broken argument is not a result. It is a coincidence with good manners.
| Incident | What the matched control reported | What it could not distinguish |
|---|---|---|
| Right-tail occupancy | Controls at 7 of the top 10, base rate 3.5 of 10; control median above treatment median in 3 of 3 panels | A broken treatment from a control that was never null |
| Sector placebo | Placebo and treatment passing on the same evidence, at the same edge | The claimed mechanism from ordinary sector beta |
| Mandates in the null arm | Books labelled null producing outcomes no null should produce | A null from an unlabelled enforced mechanism |
§3
The fix is an absolute benchmark: series that reproduce a market's real moments and contain nothing to find. Two generators, because there are two jobs.
| Generator | Construction | What it buys | Status |
|---|---|---|---|
| Edge-free null | Regime-switching process fitted to a market's real moments — unconditional volatility, volatility-clustering persistence, tail index, cross-asset correlation structure, the trading calendar — sampled so that no exploitable relationship exists in the output | A true false-positive rate: how often the pipeline invents a survivor when there is nothing to find | Built. Out-of-sample moment-match table owed |
| Planted mechanism | The same calibrated series with a known relationship injected — a cointegrating vector at a specified half-life, band width and noise; a bounded process with a known post-limit drift; an episodic intervention of specified width and frequency | A power curve: the smallest dislocation the search could have found, swept from zero upward | First curves landed. Under re-audit |
The second generator matters more than the first, and we had never built it. Every kill verdict this house has issued carries an unknown false-negative rate. A gate whose power curve begins above the dislocations that real obligations actually pay is not conservative — it is blind, and its kills are statements about the search rather than about the market. So a detection floor is now published with every kill, and a kill without one is not a finding. The floor values themselves are proprietary; the requirement and the direction are published.
The first generator's headline result is uncomfortable and belongs in public: run at full search width on panels containing nothing to find, the pipeline produced a best-of-search survivor on every panel. Search width is the main driver — a deliberately pinned narrow search cut the manufactured share by an order of magnitude. That share, and the two configuration counts behind it, are withheld as search-width calibration; the ratio between them is the part that generalises. The evidence on width is not unanimous and we will not present it as though it were: a second narrowing, run on a different panel family, roughly halved the manufactured bar without collapsing it and left power against the reference distribution at approximately zero, which reads as selection-limited rather than search-size-limited. The reading matters and we state it here rather than three paragraphs later: manufacture happens upstream of the gate, at optimisation, on a single panel, before the specificity and forward legs run. Those downstream legs are what remove these. The cost being reported is optimisation, gate, forward and review capacity consumed by candidates that were never real — not the validity of verdicts issued after the full sequence. The full treatment of that measurement is We measured our own instrument (2026-08).
Turning the instrument on our own survivors is the part we would rather not print. Nineteen books that had already cleared the matched-control gate were re-scored against a null-winner distribution generated for each book's own universe and window. Two cleared. Three were marginal. Fourteen sat below (Fig. 2).
§3, continued
The surrogate-data literature is explicit that a surrogate must preserve the nuisance structure and destroy only the hypothesised signal. Our corpus holds 24 papers on surrogate-data and block or stationary bootstrap construction, and 32 on data-snooping, reality-check and deflated-Sharpe corrections; they are counted across several shelves rather than sitting in one, and they are listed in reading. That requirement is the whole design constraint, and it cuts both ways: a generator too stationary, with the wrong correlation dynamics or a missing stylised fact, is not a null either. A survivor on such data proves nothing about real data, and a failure to survive proves nothing either. The generator simply becomes a new place for a silent zero to live.
So the house condition, written into the charter before the first batch ran: a generator that cannot pass its own moment-matching test is not a null, and any verdict resting on it is void. The moment-match table — volatility-clustering persistence, tail index, cross-sectional dispersion, autocorrelation at the horizons we trade, measured out-of-sample against the real series — is published alongside every false-positive rate. Where it is missing, the batch is void rather than probably fine.
As of this writing the table is owed for the first batch. The batch is therefore void by our own rule, the campaign built on it is under re-audit, and the standing instruction for that review is that where an artefact cannot be located or re-run the claim is void, not probably fine. We are publishing a false-positive measurement while declaring that the instrument producing it has not yet passed its own test. That is the honest state of the record.
The gate leg built on the generator is registered accordingly. It is born-informational: computed and reported in every verdict, and its failure blocks nothing. The flip to blocking is pre-registered and waits on proof. Verification so far is narrow and we will not oversell it — on a book with a documented answer the leg reproduced 3 of 3 verdicts, and higher-replication nulls landed on the documented bars. Untrusted generator states surface as loud statuses and never as a pass.
§4
In early July we wrote a section headed what we deliberately did not build, and this was in it:
By the last day of the same month, the charter for the next campaign said this:
Both sentences are correct. The ban was right on its own terms. A fitted-path generator cannot contain the structure a strategy exploits, so using synthetic paths to validate an edge flatters whatever the generator cannot represent — and we had already spent real days trading a geometric-Brownian ghost before that lesson took.
The adoption is right on different terms. A generator that contains nothing to find is not being asked to validate an edge. It is being asked how often the pipeline invents one. The question is not "does this strategy work on synthetic data", it is "how often does our machinery manufacture a winner from noise it cannot distinguish from a market". For that question the generator's inability to contain a real edge is not a defect. It is the specification.
The rule that survives both positions: synthetic data is banned from edge validation and mandatory for false-positive measurement.
The reversal is bounded. Synthetic nulls replace the absolute benchmark we never had; they do not replace matched controls, which remain the only defence against a real-but-different mechanism explaining a result. A book clears both or it holds a partial verdict.
We publish the reversal rather than quietly updating the position, for the same reason we publish refutations. A house that shows only its current doctrine is showing you its selection bias.
§5 — the standing rule
Cross-references are hand-maintained. There is no build step and no generated backlink; a link check runs before publish.
2026-08 — first published. Items in this library are not silently edited. A wrong item is corrected by a new item that cites it by title and date, and the status line above carries the result. A retracted item keeps its title, its date and its place in the index: withdrawal is a status, not a deletion.
Withheld from this paper, deliberately and by policy: the p-value attaching to the sector placebo and to the panels in Fig. 1, because the book set is live-adjacent; the null distributions and their per-venue levels; detection floors in absolute units; the search widths at which the manufactured-survivor rates were measured; the fitted parameters of both generators; the bar intervals and holding horizons our stores and searches run at; and the identities of every book, instrument, venue and market named in the source records. Mechanism classes are published, including the frequency an obligation itself specifies. The markets we screened are not, and neither are the horizons we screened them at — a list of where and at what scale we looked stays live information even when every individual result in it is dead.
A declared redaction is stronger than an omission. It tells you the number exists, that we hold it, and that we chose not to print it.