AXOQUANT

Library

2026-08 · working paper · current

Machines propose. Statistics dispose.

We do not claim capability from machine learning. Models author strategy code and generate hypotheses at volume in this shop, and nothing a model emits is permitted to adjudicate. The gate is arithmetic.

What follows is an architecture description and three measurements taken against ourselves: that our own language-model judges herd, that a learned scorer out-ranks the hand-built gate we still refuse to let it replace, and that the dominant failure mode of a machine-learning estate is not a crash but a component returning success while doing nothing. Every claim here is negative or self-limiting. That is deliberate. A capability claim invites a proof we cannot give without publishing the thing that pays us.

§1

The division of labour

Ideas enter as queued experiments — from scouts reading outside literature, from a regulatory and rulebook feed, and from the system's own follow-up proposals off contradictions in its record. A novelty gate decides whether a proposal is a genuinely new mechanic or a re-parameterisation of one we already hold; a re-parameterisation is what the optimiser does anyway, so it never triggers authoring. What survives that is written as a strategy module by a tiered cascade of models. Every authored module, from any tier, faces the same admission test: it must import cleanly and fire trades on real bars. Code that does not fire is not a strategy, and it is discarded before it costs a single optimisation hour.

From that point on, no model opinion enters any verdict. The claim is verifiable by structure rather than by benchmark: no model output is an input to any leg of any gate, and that is checkable by reading the gate. It is a governance property, not a performance one, and it is the opposite of a capability boast — a firm that has wired a model into its decisions cannot make it.

Division of authority in the research pipeline, as at 2026-08. No entry in the right column takes a model output as an input.
Machines proposeStatistics dispose
Hypothesis and mechanic generationSurvivability floors on the bootstrapped return path
Strategy code authoringSpecificity p-values against matched control ensembles
Novelty screening against the existing bookAttribution of profit to the named mechanism
Naming, summarising, triage of the recordLeverage taken as the worse of bootstrap and engine margin
Quality and duplication flags on written outputPromotion, retirement and every kill verdict

The scarce resource is stated honestly. It is not compute. It is the supply of genuinely new mechanics, and the tokens spent writing code for them. A sweep searches the parameters of one mechanic; it cannot invent a second one. A busy cluster is not evidence of progress.

§2

We measured our own judges, and they herd

Language models are widely used as reviewers, and the assumption underneath that practice — that a second model is a second opinion — is testable. We tested it on our own tooling twice.

The first test asked two judges of different lineages to call survive or fail on high-scoring candidates from in-sample information alone, with the forward outcome held back. Both rejected everything. They agreed with each other completely and trivially, and scored exactly the base rate. There is no diversity to be had between two reviewers when neither has any discrimination to disagree about.

The second was an instrument built for the purpose, scoring paired judge verdicts against ground-truth forward outcomes. It returned three findings in the same direction: the judges' errors are strongly correlated, so they fail together; their discrimination against survival is indistinguishable from chance; and the second judge adds essentially nothing over the first. The direction and the sample are published here; the correlation and discrimination values for our own review instrument are withheld, as calibration (see Redaction, below).

FIG. 1 — TWO JUDGES, ONE OPINION
unanimous 95% both wrong 46% split verdict 5%
Paired verdicts from two judges of different model lineages, scored against ground-truth forward outcomes, n = 82 pairs. The judges were unanimous on 95% of pairs and both wrong on 46% of all pairs; they split on 5%. Unanimity is therefore not evidence of correctness — it is the signature of a herd. Their discrimination against survival was indistinguishable from chance on this sample; the value for our own instrument is withheld. Consequence: judges were demoted to advisory duty and consolidation moved into code. No mark in this figure is orange, because nothing in it survived a gate.

§2.1

What we changed, rather than what we concluded

The result is only worth publishing because it had a consequence. Judges no longer touch survival verdicts; they do novelty, quality and duplication work, where their output is checkable by a human in seconds. Where a panel is still used, it is built against its own tendency to converge.

The public literature reached the same place before we did, which is why our reading list carries a shelf on crowds, herding and correlated error. We cite it rather than claim the finding as ours; what is ours is the measurement on our own instrument, and the demotion that followed.

§3

A learned scorer beat our hand-built gate

Our survivability gate is a conjunction of thresholds over a fixed set of in-sample legs. We asked whether a model over exactly the same legs — no new information, no new features — predicts forward outcomes better than the thresholds do. It does, and the ordering is stable across two independent bake-offs.

The ordering we publish: a tabular foundation model ranks first; pruned tree ensembles come next; our threshold conjunction and a logistic regression over the same legs tie at the bottom. The mechanism is in that last fact. A linear model over the legs only matches the gate, so the lift is not information the gate lacked — it is nonlinear interaction between legs: a weak reading on one leg is tolerable conditional on a strong reading of another, which an AND-of-thresholds cannot represent by construction. Which legs interact, and in which direction, is withheld as calibration. The gate is not wrong. It is the wrong shape.

The measurements: grouped cross-validation by strategy on n = 685 forward-evaluated cells, replicated in a full bake-off at n = 708–764 with every candidate scored against a 50-times shuffled-label null, and checked walk-forward by training on one month and testing on the next. Every candidate cleared the null; the discrimination values, and the margin between them, are withheld as calibration of our own instrument.

Then the part that is the actual content of this section — the policy.

A self-denying ordinance published in full is worth more than an accuracy number, and it is the only part of this section a reader can hold us to. Related reading: our shelves on tabular deep learning and evaluation metrics.

§4

A reversal, published

That programme had already been declared dead. On an early pilot we concluded the legs were exhausted and carried no learnable signal, and we stated it twice. The conclusion was wrong, for two nameable reasons: the pilot was small, and a defect in the feature pipeline was nulling one of the inputs, so a leg that existed was being measured as if it did not.

The standing lesson is not "we were wrong once". It is that dramatic results on this material — dramatically good or dramatically dead — are almost always small-sample or leakage, and must be verified before they are believed. That applies to our own scepticism as forcefully as to a strategy that looks like a discovery. The re-measurement was therefore run against a shuffled-label null before it was accepted, which is the same treatment any candidate edge gets here.

FIG. 2 — THE SAMPLE THAT REVERSED THE VERDICT
pilot 95 cells re-measured 685 cells
The same question, the same legs, opposite conclusions. The 95-cell pilot said the legs were exhausted — a conclusion we published internally twice — and it was confounded by a small sample and a defect that nulled one feature. The 685-cell re-measurement, grouped by strategy so no strategy appears in both folds, reversed it and cleared a shuffled-label null; it is marked as the result that survived its control. The ordering it produced is published in §3; the discrimination values are withheld.

§5

The decommission

On 2026-07-11 we retired a pre-trained time-series forecasting programme, having falsified all three of the uses we had built for it — as an entry gate, as a per-trade direction classifier, and as an input to regime detection — across five days of testing. One mechanism explains every failure: the model sets its forecast band from the realised volatility of the context it is fed, so it reprices the recent past rather than anticipating the future. Its lead-lag profile against realised volatility peaks at zero, measured over 607 four-hour bars. A quantity that tracks the past coincidentally cannot lead it.

The server was stopped and the scheduled prefetch removed the same day, the hardware was reclaimed, and the one branch we did not falsify — a purpose-built fine-tune judged on incremental information over the backward volatility feature — was left open and unbuilt, because nobody has built it. A published retirement with a date on it is rarer than a published launch, and it is better evidence of how a research process actually runs. Full record: the decommission notice A forecasting programme, retired (2026-07).

§6

Silent-zero is the dominant failure mode

The characteristic failure of a machine-learning estate is not a crash. A crash is loud, dated and fixed the same day. The characteristic failure is a component that returns success while doing nothing — a health endpoint answering "ok" for a service that has lost its accelerators, a scheduler reporting a finished job that never ran, a feed returning a valid empty response. Every log line looks normal. The output is plausible. Nothing alarms.

We now treat this as a class rather than a run of bad luck, because in the week to 2026-08-02 a routine operations review turned up six instances at once, in six different subsystems.

  1. A model served from CPU for the better part of a day after its accelerators were physically removed from the machine, answering its health endpoint ok throughout.
  2. A reasoning model given too small a token budget spent the entire budget thinking and returned empty content, contributing zero for a week while every log line looked normal — and collapsing a deliberately multi-lineage review panel to a single lineage without anyone noticing.
  3. A scheduled ingest failed on missing accelerators while the service manager logged finished successfully, because the command lines carried a prefix that suppresses the exit code.
  4. An archiver died inside its own error-logging path — the fail-loud loop killed itself — and then lay dead for days under a supervisor that only started it at boot.
  5. A scheduler daemon was dead for most of a day, holding its cores out of the pool, while a wake timer sent packets every ten minutes to a machine that was already awake.
  6. A work queue was empty. Not because the work was done, because both of its producers had died. An empty queue is a fault signal here now.

None of these corrupted a verdict on its own. Together they establish the risk class, and they are the reason the engineering is built to fail loudly rather than plausibly.

FIG. 3 — HOURS OF SUCCESSFUL-LOOKING NOTHING
empty output 168 h dead archiver 84 h no accelerator 19 h dead scheduler 18.5 h
The four instances of the six with a measurable undetected duration, from one operations review in the week to 2026-08-02. Each component reported success, health, or nothing at all while producing no output. The longest ran 168 hours. Marks are grey throughout: nothing here survived anything. The remaining two instances — a scheduled job logged as finished that never ran, and a queue empty because its producers had died — have no meaningful duration, only a discovery date.

§6.1

Fail loud, not plausible

Given a choice between a component that breaks visibly and an approximation that is wrong invisibly, we pay for the first. The clean case is our own execution layer: when it broke, the books submitted 700 orders over a twelve-hour parallel window and filled none of them. Every one was denied or rejected at the risk and matching layers. The defect was undeniable, diagnosed and fixed the same day, and the honest conclusion recorded at the time was that the comparison the run was meant to settle could not be made at all. A vectorised approximation of the same books would have returned a confident, well-formed, entirely fictitious result.

The operational form of the doctrine travels further than the anecdote. Every data source carries a declared minimum expected item count, so a successful empty response is an alert rather than a no-op. A source that is legitimately empty and one that is silently empty are different states and are recorded as different states. Success with zero rows is an error. A plausible empty result is never allowed to pass as a finding.

§6.2

A metric's calibration belongs to the pipeline that produced it

One more failure worth publishing, because it is the most transferable thing on this page and it costs us nothing to give away. In an earlier era of our search, one gate leg was the single best out-of-sample predictor we had. We then corrected the trial budget underneath the search — the same leg, computed the same way, on roughly 216 graded cells at the corrected budget — and it inverted, from strongly predictive to worse than a coin flip. The leg is not named and the two discrimination values are withheld; the direction of the flip is the publishable part.

In the earlier era the metric had been measuring an artefact of an under-powered search. When the search got the budget it needed, the artefact went away and took the apparent signal with it. The rule that follows is standing: a metric's calibration is a property of the pipeline that produced it, not of the metric alone. Re-check every leg's discrimination whenever the trial budget or the search space moves, and treat any leg validated under a mis-specified budget as suspect until re-validated.

§7

Architecture, not capability

Inference runs local-first on hardware we own. The panel is deliberately multi-lineage, so cross-review is not one model agreeing with itself — and §6 records what happens when a lineage silently drops out. Neither fact is a claim about quality. They are structural properties, chosen so that the failure modes we care about are the ones we can see.

The authoring inversion is the part of this we would defend hardest, and it runs against the trend. We do not ask the author model to be clever. We ask it to remember. Three memories, keyed by strategy family, go into every authoring call: the code of strategies that passed the gate, as exemplars; a registry of mechanics that were tried and died, with the reason, as anti-examples; and semantic retrieval over our own research corpus, so the author sees what we already know about the kind of edge it is being asked to build.

The evidence is an A/B against an ungrounded control on the same model, so the only variable was memory: grounding cut the rate of authored strategies that never trade by roughly twentyfold. The two rates themselves are withheld. The model did not change. The memory did.

And the honest frontier, which is not the one a capability story would choose. A bigger author is mostly the wrong move. Of the ideas we discard, essentially none fail on code — they fail on logic, as untradeable or seldom-firing mechanics, or they are correctly rejected as duplicates. The number of survivors is bounded by how many real edges the market offers, not by the author's fluency. What compounds is the corpus of our own survivors and corpses, which is not reproducible by anyone else, and which exists only because we kept the failures. Half of intelligence is learning from failure, and most systems throw that half away.

The boundary, stated plainly: this is a description of what we built. It is not a claim about what it can find. We publish no benchmark of our own stack.

§8

What we will not claim

The reason is one sentence. Any capability claim invites a proof we cannot give without leaking, and a claim that cannot be falsified is marketing.

Stands on

Cited by

The decommission notice A forecasting programme, retired (2026-07) cites §5 of this paper as the architecture's one published retirement. Cross-references are hand-maintained; a link check runs before publish.

Revision history

2026-08 — first publication, status current. Title and date are fixed at publication and are how this item is cited. It will not be silently edited; a correction is issued as a new item that cites this one.

Redaction

Withheld from this paper, deliberately and by standing policy: the correlation and discrimination values for our own judges and scorers; the identity of the gate leg that inverted in §6.2 and its two values; the trial budgets, thresholds and floors of the gate; model identities, sizes and serving locations; the size of our survivor and failure corpora; and the contents of the mined family recipes, which are distilled tradeable parameters. Where a value is withheld, the sentence says so.