Reading list
This is the reference set our house papers cite and the set we check before reaching for a default. It is curated, not comprehensive. Nothing here was added because it is famous, and inclusion is not endorsement — most of what is on these shelves was read for framing and never built.
Shelf sizes are a map of where our problems were, not of the literature's importance. A bibliography is decoration. A bibliography with a verdict against each entry is a record of what we actually did, which is the only version worth publishing.
Every entry listed below is public literature and none of it is ours. Each shelf prints a selection — 163 entries across the twenty — with authors, year, venue and a link that was fetched and checked. The count in a shelf heading is that shelf's size in our store and is not the length of the list under it; the two are counted over different populations, and the mismatch is set out under how to read an entry rather than reconciled away.
Five terms, closed set, defined once. A verdict is assigned from our own run record rather than from our opinion of the paper, and prints only where that record holds one — 10 of the entries below, all on the two featured shelves. An entry with no verdict has not been through an assignment pass, which is not the same as having been read and dismissed.
Most assigned entries are Background, and publishing that ratio is the point. A reading list on which everything was adopted is a reading list nobody read; it is a list of citations assembled after the decisions were already made. The distribution is drawn below rather than asserted.
An entry prints authors, year, title, venue, and one line on what the paper settles or fails to settle for us. Authors and years come from our own store, the arXiv API or the Crossref API, and from nowhere else — not from recall, not from inference off a title. Where a field did not resolve it is omitted rather than guessed or filled with a placeholder. Order within a shelf is the curator's rather than alphabetical: the anchor results first, then the recent work that argues with them.
Links were fetched, not assumed. Of the 163, 113 are arXiv links that returned 200; the remaining 50 are DOIs, of which 18 resolved through to a publisher page for an automated client and 32 resolved but were then refused by the publisher, which is a bot policy rather than a broken link — those open normally in a browser. A paper with no verifiable link would print without one; on this pass every listed entry had one.
A verdict prints only where our run record holds one for that entry, which is 10 of the 163. Assignment is complete on the two featured shelves as those shelves stood in the store, so a featured shelf can still print an untagged entry: it was resolved from the wider corpus after that pass and has not been through it. Used in appears only where a house item genuinely leans on the paper. Both are sparse on purpose.
The heading count and the list length answer different questions. Fig. 1 counts 244 files in the curated store on 2026-08-02. The entries are resolved from the paper corpus behind that store — 913 rows typed as paper, 802 after cleaning, 163 of which resolved to a citation and a link we could check. Shelf membership is assigned separately in each. The list is shorter than the heading on nine shelves, equal on three, and longer on eight. Neither number is wrong and neither is a correction of the other; they count different things, and the honest move is to print both and say so.
Featured shelf
The public literature on machine agents trading, judging and negotiating in markets, together with the older wisdom-of-crowds and diverse-problem-solver results it rests on. We read this shelf because we run such agents: models author strategy code and generate hypotheses at volume across our estate. Naming it costs us nothing, because it is public science and none of it is ours.
The shelf splits. One half reports machine agents performing; the other reports multi-agent teams holding experts back, information leaking into apparent profit, and valid signals failing at regime boundaries. Our own measurement agrees with the sceptical half, and we ran it rather than asserted it.
Two judges from different model lineages were given the same high-scoring candidates and asked to call survive or fail from in-sample information alone. Both rejected everything. They agreed with each other completely and trivially, and they scored exactly the base rate of the sample they were shown. A second model adds no diversity when there is no discrimination to disagree about. The judges were demoted to advisory and consolidation moved into code — nothing a model emits adjudicates anything here. The accuracy figure maps directly onto our own corpus base rate and is withheld; the direction, the agreement and the policy consequence are published. The argument in full is in machines.
11 entries listed · shelf of 27 files in Fig. 1
Hong, Page. 2004. Groups of diverse problem solvers can outperform groups of high-ability problem solvers. Proceedings of the National Academy of Sciences.
The formal case that diversity beats individual ability. Our two judges had neither and scored the base rate.
Bikhchandani, Sharma. 2000. Herd Behavior in Financial Markets. IMF Staff Papers.
Separates intentional herding from clustering that only looks like it. The distinction a copy-flow strategy stands or falls on.
Barber, Odean. 2000. Trading Is Hazardous to Your Wealth: The Common Stock Investment Performance of Individual Investors. The Journal of Finance.
Account-level records rather than survey answers. Turnover predicts underperformance.
Simoiu, Sumanth, Mysore et al. 2019. Studying the “Wisdom of Crowds” at Scale. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing.
Crowd accuracy measured across domains instead of asserted from one.
Pappu, El, Cao et al. 2026. Multi-Agent Teams Hold Experts Back. arXiv preprint.
Adding agents removed accuracy the strongest member already had. Our own judge measurement points the same way.
Henning, Ojha, Spoon et al. 2025. LLM Agents Do Not Replicate Human Market Traders: Evidence From Experimental Finance. arXiv preprint.
Experimental-finance baseline the agents fail to reach. Read against every paper on this shelf that reports agents performing.
Fatouros, Metaxas. 2026. Signal or Noise in Multi-Agent LLM-based Stock Recommendations? arXiv preprint.
Asks whether the ensemble adds information over one model. The question we ran in-house and answered no.
Dudley, Magdaleno. 2026. Prediction Markets Underperform Simple Baselines For Infectious Disease Forecasting. arXiv preprint.
A crowd mechanism losing to a naive baseline outside its home domain.
Xu, Li, Liu et al. 2026. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias. arXiv preprint.
Traces LLM-as-judge bias to mechanism rather than prompt. Read after we demoted our judges to advisory.
Xiao, Sun, Di Luo et al. 2024. TradingAgents: Multi-Agents LLM Financial Trading Framework. arXiv preprint.
The multi-agent trading architecture stated by its authors. Held to argue with, not to adopt.
Agrawal, Teo, Vazquez et al. 2025. Evaluating LLM Agent Collusion in Double Auctions. arXiv preprint.
Collusion as an emergent property of the market, not of the model.
One Adopted on this shelf. The diverse-problem-solver result is why our review runs on models of deliberately different lineage rather than on more copies of one — and why, when we measured the ensemble and found no discrimination, we did not reach for a third model.
Featured shelf
Mostly the time-series foundation-model literature of the last four years — pretrained probabilistic forecasters, the transformer architectures underneath them, and the benchmarks that evaluate them. Individual papers are not named in this paragraph: every entry listed below carries its authors, year and a checked link, and the shelf is larger than the list. We implemented from this shelf, pre-registered what would falsify the implementation, and decommissioned the programme on a dated day.
The deployed model was tested on three separate trading uses across five days and failed all three. As an entry gate it added no win-rate edge at honest trade counts, and the high-scoring configurations were artifacts of nineteen to twenty-eight trades, rejected by a minimum trade count written down before the run. A tabular classifier on the same causal features cleared no out-of-sample bar. And neither track improved the market-regime detector. The floors were pre-registered; their values are withheld, as all our cut-offs are.
One mechanism explains all three. The model sets its forecast band from the realised volatility of the context it is fed, so it reprices the recent past rather than anticipating the future. Measured over 607 four-hour bars, its predicted band width correlated with trailing realised volatility at +0.82, its correlation with forward volatility sat below the plain backward baseline, and its lead-lag profile peaked at zero to one day forward. A quantity that tracks the past coincidentally cannot lead it. The honest caveat, published with the verdict: the label overlap was 84 to 88 days over a single episode that ran mostly in one direction. Small sample — but the same-feature correlation and the lag-zero peak are not marginal, and they are the exact failure the mechanism predicts. The forecast server was stopped and disabled on 2026-07-11.
The distinction a careless reading list destroys: this is a negative about our use, not about the papers. They are competent work on a problem we could not make pay at our horizon, and the one branch left unfalsified — a fine-tune trained on forward realised volatility rather than next-bar price, judged on incremental correlation after the backward-volatility feature — has not been built by anyone here. The account in full is in decommission.
11 entries listed · shelf of 21 files in Fig. 1
Ansari, Stella, Turkmen et al. 2024. Chronos: Learning the Language of Time Series. arXiv preprint.
The model we deployed and decommissioned. Forecast trailed volatility rather than leading it.
Das, Kong, Sen et al. 2023. A decoder-only foundation model for time-series forecasting. arXiv preprint.
Zero-shot forecasting as a pre-training problem. The premise our decommission tested.
Woo, Liu, Kumar et al. 2024. Unified Training of Universal Time Series Forecasting Transformers. arXiv preprint.
Cross-domain pre-training for a single forecaster.
Goswami, Szafer, Choudhry et al. 2024. MOMENT: A Family of Open Time-series Foundation Models. arXiv preprint.
Open time-series foundation models, released with their evaluation.
Gruver, Finzi, Qiu et al. 2023. Large Language Models Are Zero-Shot Time Series Forecasters. arXiv preprint.
Numeric forecasting from a text model. Sets the claim; the finance evidence is elsewhere.
Nie, Nguyen, Sinthong et al. 2022. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. arXiv preprint.
Patching, and the channel-independence choice that most later architectures inherit.
Liu, Hu, Zhang et al. 2023. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. arXiv preprint.
Inverts the attention axis. Held because it changes what the model treats as a token.
Huang, Xu, Darlow. 2026. How Good Can Linear Models Be for Time-Series Forecasting? arXiv preprint.
The control that most deep forecasting papers owe and few run.
Aksu, Woo, Liu et al. 2024. GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation. arXiv preprint.
A benchmark rather than a model. Read for the evaluation protocol, not the leaderboard.
Brini. 2026. Forecasting Realized Volatility with Time Series Foundation Models: A Comparison with Econometric Benchmarks. arXiv preprint.
Foundation models against econometric benchmarks on the one target where the econometrics is strong.
Barunik, Hronec, Tobek. 2024. Forecasting stock return distributions around the globe with quantile neural networks. arXiv preprint.
Quantile networks on the distribution rather than the mean.
Nine of the twenty-one are Read, not pursued, and the reason is the same for all nine: one model from this family was implemented and the mechanism that killed it is a property of the family, not of the implementation. We did not run eight more experiments to confirm a mechanism we had already identified. That is a stated reason, not a verdict — if the mechanism argument is wrong, these nine are owed a test.
Structure, counts, the problem that sent us there, and the entries. These eighteen have had no verdict pass, so they print citations without verdicts rather than print a verdict we have not earned. Each shelf carries an anchor so a house paper can deep-link a single citation.
Microstructure41
Fill models, queue position, order-book dynamics, impact and execution under latency. The largest shelf because execution is where our verdicts die: vectorised screens overstate edge, and our own fills ledger has voided our own results more than once.
13 entries listed · shelf of 41 files in Fig. 1
Amihud. 2002. Illiquidity and stock returns: cross-section and time-series effects. Journal of Financial Markets.
The illiquidity measure most later work builds on, cross-section and time-series.
Obizhaeva, Wang. 2013. Optimal trading strategy and supply/demand dynamics. Journal of Financial Markets.
Impact with a resilience term. Our execution assumptions are a special case of this.
Budish, Cramton, Shim. 2015. The High-Frequency Trading Arms Race: Frequent Batch Auctions as a Market Design Response. The Quarterly Journal of Economics.
Names the mechanism as a market-design defect and proposes the fix. A mechanism paper, not a pattern paper.
Menkveld. 2013. High frequency trading and the new market makers. Journal of Financial Markets.
What the fast intermediary actually does, from a venue's own records.
Foucault, Hombert, Roşu. 2016. News Trading and Speed. The Journal of Finance.
Separates speed applied to news from speed applied to order flow.
Brandt, Kavajecz. 2004. Price Discovery in the U.S. Treasury Market: The Impact of Orderflow and Liquidity on the Yield Curve. The Journal of Finance.
Order flow priced into the yield curve, measured rather than assumed.
Cont, Cucuringu, Zhang. 2021. Cross-Impact of Order Flow Imbalance in Equity Markets. arXiv preprint.
Whether one asset's imbalance moves another. The multi-asset case our single-asset impact model ignores.
Ding, Hanna, Hendershott. 2014. How Slow Is the NBBO? A Comparison with Direct Exchange Feeds. Financial Review.
Consolidated feed against direct feeds. The latency our backtests silently assume away.
Briola, Bartolucci, Aste. 2024. Deep Limit Order Book Forecasting. arXiv preprint.
Horizon and stationarity treated as first-class. Read for the evaluation, not the model.
Prata, Masi, Berti et al. 2023. LOB-Based Deep Learning Models for Stock Price Trend Prediction: A Benchmark Study. arXiv preprint.
Benchmarks the family against itself. The reported gaps shrink under a common protocol.
Mesfin. 2026. Structural Limits of OHLCV-Based Intraday Signals in MNQ Futures: A Systematic Falsification Study. arXiv preprint.
A falsification study on the bar data we use. Publishes a negative on its own hypothesis.
MacKenzie. 2018. Material Signals: A Historical Sociology of High-Frequency Trading. American Journal of Sociology.
A sociology of the trade, not a model of it. Held because it describes where the fills come from.
Aloosh, Li. 2024. Direct Evidence of Bitcoin Wash Trading. Management Science.
Volume that is not volume. The data-quality problem before the trading one.
Volatility33
Realised and implied estimation, stress and uncertainty indices, tail dependence. Volatility is the input to sizing, to regime labels and to every calibrated null we fit, so the shelf is about measuring it honestly rather than predicting it.
11 entries listed · shelf of 33 files in Fig. 1
French, Schwert, Stambaugh. 1987. Expected stock returns and volatility. Journal of Financial Economics.
The early statement of the relation, and the estimation problems that come with it.
Ang, Hodrick, Xing et al. 2009. High idiosyncratic volatility and low returns: International and further U.S. evidence. Journal of Financial Economics.
International evidence on the puzzle. Read because our sizing rule leans the opposite way to the anomaly.
Christensen, Podolskij. 2026. Realized range-based estimation of integrated variance. arXiv preprint.
Range estimators against realised variance. The efficiency argument for using bars we already store.
Christensen, Christiansen, Posselt. 2020. The economic value of VIX ETPs. Journal of Empirical Finance.
Whether the instrument delivers the exposure it names.
Cipollini, Cruciani, Gallo et al. 2026. VOLatility Archive for Realized Estimates (VOLARE). arXiv preprint.
An archive of realised estimates. Held for reproducibility of the input, not for a method.
Liu, Fu, Hong. 2025. Forecasting realized volatility in the stock market: a path-dependent perspective. arXiv preprint.
Path dependence in the volatility process rather than a longer lag.
Goel, Pasricha, Magris et al. 2025. Foundation Time-Series AI Model for Realized Volatility Forecasting. arXiv preprint.
Foundation models on the volatility target. Read alongside our own decommission note.
Wade. 2026. Do Better Volatility Forecasts Lead to Better Portfolios? Evidence from Graph Neural Networks. arXiv preprint.
Asks whether the forecast improvement survives into the allocation. Most volatility papers do not ask.
Bi, Calhoun. 2026. Conditioning on a Volatility Proxy Compresses the Apparent Timescale of Collective Market Correlation. arXiv preprint.
Conditioning on a proxy changes the measured timescale. A warning about our own regime labels.
Christensen, Liu, Liu et al. 2026. The realized copula of volatility. arXiv preprint.
Dependence between volatilities rather than between returns.
Maghyereh, Awartani, Bouri. 2016. The directional volatility connectedness between crude oil and equity markets: New evidence from implied volatility indexes. Energy Economics.
Spillover measured with a direction attached.
Regime18
Regime-switching and hidden-Markov dynamics, changepoint detection, state-space estimation. Read first for labelling exposure, then re-read for a different job entirely: our edge-free generator is specified as a regime-switching model fitted to a market's real moments. Specified, not yet demonstrated — the moment-match table that would show the fit holds is owed, and until it lands the generator is a design rather than a validated null.
10 entries listed · shelf of 18 files in Fig. 1
Oelschläger, Adam. 2020. Detecting bearish and bullish markets in financial time series using hierarchical hidden Markov models. arXiv preprint.
Hierarchical hidden Markov states fitted to price alone. The baseline our labels have to beat.
Shu, Yu, Mulvey. 2024. Downside Risk Reduction Using Regime-Switching Signals: A Statistical Jump Model Approach. arXiv preprint.
A jump model used as an exposure signal rather than a description.
Zakamulin, Giner. 2024. Optimal trend-following rules in two-state regime-switching models. Journal of Asset Management.
Derives the rule from the state model instead of fitting the rule and naming the states after it.
Tsaknaki, Lillo, Mazzarisi. 2024. Bayesian Autoregressive Online Change-Point Detection with Time-Varying Parameters. arXiv preprint.
Online detection with time-varying parameters. Online is the constraint that rules most of this literature out for us.
Khamesi, Cohen, Adams et al. 2026. CHASM: Online Changepoint Detection in Temporal and Cross-Variable Dependence. arXiv preprint.
Changepoints in cross-variable dependence, not in the level.
Orton, Gebbie. 2024. Representation Learning for Regime detection in Block Hierarchical Financial Markets. arXiv preprint.
Learned states in a block-hierarchical market. Read for whether learned states are causal.
Hammond. 2026. Geometric Observables for Financial Regime Detection. arXiv preprint.
Regime as geometry of the path.
Yi, Mehra, Chen et al. 2026. Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market. arXiv preprint.
Text as a regime input on the Treasury market.
Thumm. 2025. Causal Regime Detection in Energy Markets With Augmented Time Series Structural Causal Models. arXiv preprint.
Structural causal models on a market with a physical constraint underneath.
Taljaard. 2025. Regime-based Portfolio Optimisation: A Hidden Markov Model Approach for Fixed Income Portfolios. SSRN Electronic Journal.
Hidden Markov states used for allocation in fixed income.
Cointegration18
Long-run relationships, error correction, Hurst and Granger machinery. Read because a pair that cointegrates in-sample and not out of it is the most common false positive we generate, and the tests do not warn you.
8 entries listed · shelf of 18 files in Fig. 1
Gatev, Goetzmann, Rouwenhorst. 2006. Pairs Trading: Performance of a Relative-Value Arbitrage Rule. Review of Financial Studies.
The reference result. Read now for the decay of the reported edge as the rule became public.
Serban. 2010. Combining mean reversion and momentum trading strategies in foreign exchange markets. Journal of Banking & Finance.
Two signals with opposite signs on the same series.
Epstein, Wang, Choi et al. 2025. Attention Factors for Statistical Arbitrage. arXiv preprint.
Learned factors in place of a fixed spread definition.
Reichold, Schneider. 2025. Beyond the Oracle Property: Adaptive LASSO in Cointegrating Regressions with Local-to-Unity Regressors. arXiv preprint.
Adaptive LASSO where the regressor is local-to-unity. The near-unit-root case our screens hit constantly.
Li, Yan, Yao. 2025. Factor Models of Matrix-Valued Time Series: Nonstationarity and Cointegration. arXiv preprint.
Nonstationarity and cointegration in a factor structure rather than a pair.
Lopetuso, Caporin. 2026. The Cointegrated Matrix Autoregressive Model. arXiv preprint.
Extends the error-correction form to matrix-valued series.
Hong, Klabjan. 2025. Graph Learning for Foreign Exchange Rate Prediction and Statistical Arbitrage. arXiv preprint.
Cross-rate structure as a graph. The universe-selection step our pair screens do by hand.
Zournatzidou, Floros. 2023. Hurst Exponent Analysis: Evidence from Volatility Indices and the Volatility of Volatility Indices. Journal of Risk and Financial Management.
Persistence measured on volatility indices. Read because our Hurst estimates move with the window.
Risk management11
Stops, trailing exits, drawdown control and restart rules. Read hard, because every consensus-conservative default we tested empirically moved in the deployable direction, and the stop-loss bundle was the largest single case.
11 entries listed · shelf of 11 files in Fig. 1
Artzner, Delbaen, Eber et al. 1999. Coherent Measures of Risk. Mathematical Finance.
The axioms. Read because drawdown-from-peak satisfies none of them and we use it anyway on some books.
Dybvig. 1988. Inefficient Dynamic Portfolio Strategies or How to Throw Away a Million Dollars in the Stock Market. Review of Financial Studies.
Prices path-dependent strategies against their static equivalent. The cost of the path, stated as a number.
Kaminski, Lo. 2010. When Do Stop-Loss Rules Stop Losses? SSRN Electronic Journal.
The conditions under which a stop helps, and the return process it needs. Our high-win-rate books fail those conditions.
Lei, Li. 2009. The Value of Stop Loss Strategies. SSRN Electronic Journal.
Measures the stop rather than assuming it.
Han, Zhou. 2014. Taming Momentum Crashes: A Simple Stop-Loss Strategy. SSRN Electronic Journal.
A stop applied to the one strategy family with a documented crash. The case where the consensus default earns its place.
Zambelli. 2016. Determining Optimal Stop-Loss Thresholds via Bayesian Analysis of Drawdown Distributions. arXiv preprint.
Fits the threshold to the drawdown distribution instead of picking a round percentage.
Leung, Zhang. 2017. Optimal Trading with a Trailing Stop. arXiv preprint.
The trailing case solved rather than simulated.
Hsieh. 2023. On Data-Driven Drawdown Control with Restart Mechanism in Trading. arXiv preprint.
Control with a restart rule. Restart is the part most drawdown papers leave out.
Carr, Jarrow. 2008. The Stop-Loss Start-Gain Paradox and Option Valuation: A new Decomposition into Intrinsic and Time Value. Financial Derivatives Pricing.
Shows what the stop is worth as an option. Read before treating a stop as free.
Koike, Hofert, Tsunekawa. 2026. Measuring multivariate maximal tail dependence. arXiv preprint.
Dependence in the tail rather than in the covariance.
Li, Joe. 2026. Extreme Value Inference for CoVaR and Systemic Risk. arXiv preprint.
Systemic risk estimated where the data is thinnest.
Portfolio10
Construction, higher moments, hierarchical and risk-parity methods, and the ways funds fail. Read against our own evidence that de-correlating a negative-mean distribution compresses it toward its mean and destroys the right tail that carries the edge.
12 entries listed · shelf of 10 files in Fig. 1
Carhart. 1997. On Persistence in Mutual Fund Performance. The Journal of Finance.
Persistence explained by factors and costs rather than skill. The null every allocator owes.
Choueifaty, Froidure, Reynier. 2013. Properties of the most diversified portfolio. The Journal of Investment Strategies.
Diversification as an explicit objective, with its properties derived rather than asserted.
de Prado. 2018. The 10 Reasons Most Machine Learning Funds Fail. The Journal of Portfolio Management.
The failure list. We have made several of these and the paper names them better than we did.
Boyd, Johansson, Kahn et al. 2024. Markowitz Portfolio Construction at Seventy. arXiv preprint.
What survives of mean-variance once estimation error is taken seriously.
Fung, Hsieh. 2001. The Risk in Hedge Fund Strategies: Theory and Evidence from Trend Followers. Review of Financial Studies.
Trend following decomposed into option-like exposures rather than skill.
Agarwal, Daniel, Naik. 2011. Do Hedge Funds Manage Their Reported Returns? Review of Financial Studies.
Reported returns as a managed quantity. Read before trusting any published track record, including a future one of ours.
Brunnermeier, Nagel. 2004. Hedge Funds and the Technology Bubble. The Journal of Finance.
Sophisticated capital riding a mispricing rather than correcting it.
Maza. 2025. Hierarchical risk clustering versus traditional risk-based portfolios: an empirical out-of-sample comparison.
An out-of-sample comparison rather than an in-sample demonstration.
Bongiorno, Manolakis, Mantegna. 2026. End-to-end large portfolio optimization for variance minimization with neural networks through covariance cleaning. The Journal of Finance and Data Science.
Covariance cleaning inside the optimisation rather than before it.
Tasitsiomi. 2025. On the hidden costs of passive investing. arXiv preprint.
The costs the index does not print.
Alexander, Fabozzi. 2026. Measuring Strategy-Decay Risk: Minimum Regime Performance and the Durability of Systematic Investing. arXiv preprint.
Durability as minimum performance across regimes, not average performance over the sample.
Raddatz, Schmukler, Williams. 2017. International asset allocations and capital flows: The benchmark effect. Journal of International Economics.
Benchmark membership moving flows. A named, dated, mechanical trigger of the kind we look for.
Outliers9
Anomaly detection in multivariate series, conformal methods, root-cause attribution. Read for the data-quality problem before the trading one: an unadjusted corporate action looks exactly like an anomaly worth trading, and has been mistaken for one.
10 entries listed · shelf of 9 files in Fig. 1
Köhler, Heckens, Guhr. 2026. Extreme Value Analysis for Finite, Multivariate and Correlated Systems with Finance as an Example. arXiv preprint.
Extreme value theory where the sample is finite and the series are correlated, which is our case.
Uehara. 2026. Stop Suppressing the Tail: Causal Inference for Extreme Events. arXiv preprint.
Argues against the cleaning step most pipelines apply by default, ours included.
Yeh. 2026. Matrix Profile for Time-Series Anomaly Detection: A Reproducible Open-Source Benchmark on TSB-AD. arXiv preprint.
A reproducible benchmark. Held for the benchmark, not the method.
Mishra, Patil, Schockaert et al. 2026. Conditional Attribution for Root Cause Analysis in Time-Series Anomaly Detection. arXiv preprint.
Attribution after detection. Detection alone does not tell you whether it was a corporate action.
Kim, Mok, Lee et al. 2026. CANDI: Curated Test-Time Adaptation for Multivariate Time-Series Anomaly Detection Under Distribution Shift. arXiv preprint.
Test-time adaptation under distribution shift, which is the condition anomaly detectors are actually deployed in.
Hermary, Hicsonmez, Pineau et al. 2026. ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection. arXiv preprint.
Unsupervised detection with generated pseudo-anomalies, since real labels are the scarce thing.
Khosravinia, Gama, Veloso. 2026. Causally-Constrained Probabilistic Forecasting for Time-Series Anomaly Detection. arXiv preprint.
A causal constraint on the detector. The constraint we impose on every feature.
Gil, O'Donncha, Gifford et al. 2026. Adaptive Conformal Anomaly Detection with Time Series Foundation Models for Signal Monitoring. arXiv preprint.
Conformal coverage on top of a foundation model.
Jonkers, Ziegel. 2026. Generalized Conformal Predictive Systems Under Distributional Shifts. arXiv preprint.
What the coverage guarantee is worth once the distribution moves.
Caputi, Meadows. 2026. Financial Anomaly Detection for the Canadian Market. arXiv preprint.
One market, stated. Read for the data handling.
Crypto9
Perpetual funding, fee determination, equilibrium dynamics and pair mechanics. A small shelf, because the tradeable literature is much thinner than the volume of writing on the asset class suggests. Shelf size here measures the literature, not our exposure.
12 entries listed · shelf of 9 files in Fig. 1
He, Manela, Ross et al. 2022. Fundamentals of Perpetual Futures. arXiv preprint.
The no-arbitrage account of the funding mechanism. The obligation is real; whether the dislocation exceeds cost is a separate question.
Kim, Park. 2025. Designing funding rates for perpetual futures in cryptocurrency markets. arXiv preprint.
Funding as a design choice by the venue rather than a market outcome.
Zhivkov. 2026. The Two-Tiered Structure of Cryptocurrency Funding Rate Markets. Mathematics.
Structure across venues rather than a single funding series.
Chitra. 2025. Autodeleveraging: Impossibilities and Optimization. arXiv preprint.
The venue's forced-unwind mechanism, and what it cannot achieve. A named obligated party.
Krestenko, Butov, Berezovskiy et al. 2026. Dynamic Collateral Control for Permissionless Spot Perpetual Basis Trading. arXiv preprint.
The collateral leg of the basis trade, which is where the carry result usually dies.
Li, Dahmani, Cai. 2026. Arbitrage and the Stability of AMM Price Tracking. arXiv preprint.
How closely the pool tracks, and who pays for the tracking.
Urusov, Berezovskiy, Krestenko et al. 2026. Liquidity provision in CLMMs: evidence from transactions data. arXiv preprint.
Transaction-level evidence on concentrated liquidity rather than a simulation.
Baggiani, Herdegen, Sanchez-Betancourt. 2026. Competition between DEXs through Dynamic Fees. arXiv preprint.
Fee-setting as competition. Fees are the term that killed our sub-hourly lane.
Tadi, Witzany. 2023. Copula-Based Trading of Cointegrated Cryptocurrency Pairs. arXiv preprint.
Dependence beyond correlation on pairs we have screened.
Padyšák, Vojtko. 2022. Seasonality, Trend-following, and Mean reversion in Bitcoin. SSRN Electronic Journal.
Three effects on one asset, reported together rather than one at a time.
Corbet, Lucey, Urquhart et al. 2019. RETRACTED: Cryptocurrencies as a financial asset: A systematic analysis. International Review of Financial Analysis.
Retracted after publication. Kept on the shelf as the reason a citation carries its status, not just its claim.
Wu, Liu. 2026. Stability Anchors and Risk Amplifiers: Tail Spillovers Across Stablecoin Designs. arXiv preprint.
Tail spillovers across stablecoin designs. A peg is an obligation until it is not.
Tabular deep learning8
Prior-fitted networks, transformers for tables, embeddings for numerical features. Read because a learned scorer out-ranked our hand-built gate, which opened the governance question rather than settling it. It advises. It does not set a floor.
11 entries listed · shelf of 8 files in Fig. 1
Hollmann, Müller, Eggensperger et al. 2022. TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second. arXiv preprint.
Prior-fitted inference in place of fitting. The result that made a learned scorer worth testing against our gate.
Gorishniy, Rubachev, Khrulkov et al. 2021. Revisiting Deep Learning Models for Tabular Data. arXiv preprint.
The comparison that keeps returning gradient boosting to the top.
Gorishniy, Rubachev, Babenko. 2022. On Embeddings for Numerical Features in Tabular Deep Learning. arXiv preprint.
Where the gain in tabular deep learning actually comes from.
Rubachev, Kartashev, Gorishniy et al. 2024. TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning Benchmarks. arXiv preprint.
Benchmarks with temporal structure, so the leakage the standard splits hide becomes visible.
Erickson, Purucker, Tschalzev et al. 2025. TabArena: A Living Benchmark for Machine Learning on Tabular Data. arXiv preprint.
A living benchmark. Held because a static leaderboard ages badly.
Ye, Yin, Zhan et al. 2024. Revisiting Nearest Neighbor for Tabular Data: A Deep Tabular Baseline Two Decades Later. arXiv preprint.
A two-decade-old baseline still competitive. The control this field keeps skipping.
Ma, Thomas, Hosseinzadeh et al. 2024. TabDPT: Scaling Tabular Foundation Models on Real Data. arXiv preprint.
Scaling the prior-fitted approach on real data rather than synthetic priors.
Ye, Liu, Chao. 2025. A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its Capabilities. arXiv preprint.
Where the method holds and where it stops.
Deprez, Verbeke, Verdonck. 2026. Is TabPFN the Silver Bullet for Insurance Pricing? arXiv preprint.
An out-of-domain test by people with a real loss function. The answer is qualified.
Loza, Chushig-Muzo, Milara et al. 2026. Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models. arXiv preprint.
Out-of-distribution behaviour, which is the only regime we deploy in.
Baesens, Goethals, Lessmann et al. 2026. Foundation Models for Credit Risk Prediction: A Game Changer? arXiv preprint.
Another applied domain asking the same question and not settling it.
Metrics8
Ranking and discrimination measures, information coefficients, multi-class AUC. Read because which in-sample statistic predicts out-of-sample survival is a question our own corpus answers uncomfortably. That the answer is counterintuitive is published; the ranking itself is a calibration of our own screen and is withheld.
9 entries listed · shelf of 8 files in Fig. 1
Sullivan, Timmermann, White. 1999. Data‐Snooping, Technical Trading Rule Performance, and the Bootstrap. The Journal of Finance.
The universe of rules priced in, not just the winner. The correction our screens are built around.
Harvey, Liu, Zhu. 2016. … and the Cross-Section of Expected Returns. Review of Financial Studies.
Raises the t-statistic hurdle for a literature that had been reading two as significant.
Bailey, de Prado. 2014. The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. SSRN Electronic Journal.
Adjusts for selection, trial count and non-normality. Our own floor is a calibration and is withheld; the correction is not.
Bailey, Borwein, de Prado et al. 2016. The probability of backtest overfitting. The Journal of Computational Finance.
Overfitting as a probability attached to a specific backtest rather than a general warning.
Bailey, Ger, de Prado et al. 2015. Statistical Overfitting and Backtest Performance. Risk-Based and Factor Investing.
The in-sample-to-out-of-sample relation, stated as a degradation.
Waghmare, Ziegel. 2025. Proper scoring rules for estimation and forecast evaluation. arXiv preprint.
What a scoring rule has to satisfy before a comparison means anything.
Angelopoulos, Barber, Bates. 2024. Theoretical Foundations of Conformal Prediction. arXiv preprint.
Distribution-free coverage, and what it costs.
Pic, Dombry, Naveau et al. 2024. Proper Scoring Rules for Multivariate Probabilistic Forecasts based on Aggregation and Transformation. arXiv preprint.
The multivariate case, where most convenient scores stop being proper.
Wang. 2023. Calibration in Deep Learning: A Survey of the State-of-the-Art. arXiv preprint.
Confidence against frequency. A model that ranks well can still be badly calibrated, and sizing uses the calibration.
Features7
Automated generation, pruning, causal selection for multivariate series. A small shelf because most feature machinery we tried added leakage faster than it added signal, and the causal-window constraint rules out a good deal of it before we start.
8 entries listed · shelf of 7 files in Fig. 1
Zhang, Zhang, Fan et al. 2022. OpenFE: Automated Feature Generation with Expert-level Performance. arXiv preprint.
Automated generation with a pruning step. The generation is cheap; the pruning is the paper.
Sun, Li, Liu et al. 2015. Using causal discovery for feature selection in multivariate numerical time series. Machine Learning.
Selection by causal structure rather than by correlation with the label.
Penim, Pereira, Bono et al. 2026. Causal Discovery on Irregular Time Series. arXiv preprint.
Irregular sampling, which is what tick and funding data actually are.
Mazaheri, Zhang, Uhler. 2026. Relaxing Faithfulness with Intervention-Only Causal Discovery. arXiv preprint.
Weakens an assumption that market data does not satisfy.
Zheng, Verma, Gill et al. 2026. Causal Discovery in the Era of Agents. arXiv preprint.
Survey of where the field now stands. Held for scope, not for a method.
Fesanghary. 2026. Causal-TS: A Python Library for Causal Discovery in High-Dimensional and Nonstationary Time Series. arXiv preprint.
An implementation for high-dimensional nonstationary series. Read for the code, not the claim.
Babiak, Barunik, Kurka. 2026. Skewness Dispersion and Stock Market Returns. arXiv preprint.
A higher-moment feature with a stated construction.
Sidhu, Fan, Pishgar. 2026. Which Voices Move Markets? Speaker Identity and the Cross-Section of Post-Earnings Returns. arXiv preprint.
Speaker identity as a feature in post-earnings returns.
Geared funds6
Daily rebalance mechanics, compounding drag beyond volatility decay, long-horizon behaviour. A mechanism class with a named obligated party and a dated, contractual trigger, which is the shape we look for before a pattern is worth authoring.
6 entries listed · shelf of 6 files in Fig. 1
Avellaneda, Zhang. 2010. Path-Dependence of Leveraged ETF Returns. SIAM Journal on Financial Mathematics.
The compounding result. Contractual daily rebalance, dated trigger, named obligated party.
Leung, Park, Yeo. 2023. Robust Long-Term Growth Rate of Expected Utility for Leveraged ETFs. arXiv preprint.
Long-horizon growth rate derived rather than simulated.
Brown. 2023. Long-Term Returns Estimation of Leveraged Indexes and ETFs. arXiv preprint.
Estimation over horizons where the drag dominates.
Hsieh, Chang, Chen. 2025. Compounding Effects in Leveraged ETFs: Beyond the Volatility Drag Paradigm. arXiv preprint.
Argues the drag is not only volatility decay. Read against the standard explanation.
Hayase, Mizuta, Yagi. 2026. Impact of arbitrage between leveraged ETF and futures on market liquidity during market crash. arXiv preprint.
The rebalance flow hitting the futures book during a crash.
Bianchi, Goldberg. 2026. A Levered ETF Anomaly Explained. arXiv preprint.
Names a mechanism for a pattern that had been reported without one.
Pipelines4
Automated machine-learning pipelines and their operations. Read for the engineering, kept for the failure modes — the dominant one in an estate this size is a component that reports success while doing nothing, which is why ours are built to fail loudly.
2 entries listed · shelf of 4 files in Fig. 1
Shchur, Turkmen, Erickson et al. 2023. AutoGluon-TimeSeries: AutoML for Probabilistic Time Series Forecasting. arXiv preprint.
AutoML for probabilistic forecasting. Read for the engineering; the failure modes are ours to add.
Pham, Nguyen, Thi. 2026. AlgoXpert Alpha Research Framework. A Rigorous IS WFA OOS Protocol for Mitigating Overfitting in Quantitative Strategies. arXiv preprint.
An in-sample, walk-forward and out-of-sample protocol written down as a protocol.
Entropy4
Approximate and sample entropy, complexity measures for physiological and financial series. Read for complexity features. The shelf is small because the payoff was small, and we stopped rather than kept reading.
4 entries listed · shelf of 4 files in Fig. 1
Richman, Moorman. 2000. Physiological time-series analysis using approximate entropy and sample entropy. American Journal of Physiology-Heart and Circulatory Physiology.
The source paper for the estimator, from the field that validated it.
Delgado-Bonal, Marshak. 2019. Approximate Entropy and Sample Entropy: A Comprehensive Tutorial. Entropy.
The tutorial, including the parameter sensitivity that makes the measure hard to compare across studies.
Masoudi, Shahbazi, Sharifi. 2025. Complexity of Financial Time Series: Multifractal and Multiscale Entropy Analyses. arXiv preprint.
Multifractal and multiscale entropy on price series.
Giri, Kayal. 2025. Permutation extropy: a time series complexity measure. arXiv preprint.
A complementary complexity measure. Small shelf, small payoff, and we stopped.
Machine learning, general3
Scale and stability challenges, double descent. Framing for everything else on these shelves; nothing here is a method we run.
8 entries listed · shelf of 3 files in Fig. 1
Schulman, Wolski, Dhariwal et al. 2017. Proximal Policy Optimization Algorithms. arXiv preprint.
The policy-gradient method most trading RL papers use. Framing, not a method we run.
McInnes, Healy, Melville. 2018. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv preprint.
The projection we use for looking at strategy populations. It is a lens, not evidence.
Gu, Zheng, Aste. 2023. Unraveling the Enigma of Double Descent: An In-depth Analysis through the Lens of Learned Feature Space. arXiv preprint.
Double descent traced through the learned feature space.
Nikolopoulos. 2026. Spurious Predictability in Financial Machine Learning. arXiv preprint.
Predictability that appears without a mechanism. The failure mode our gate exists to catch.
Zhang, Li, Peng et al. 2026. When Alpha Disappears: A One-Switch Benchmark for Decision-Time Leakage in Financial Backtests. arXiv preprint.
A benchmark for decision-time leakage. The one-bar lag that inflated one of our own results five-fold is this failure.
Zhou, Liu, Du et al. 2026. From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems. arXiv preprint.
Determinism as a requirement in financial systems. Directly relevant to reproducing our own runs.
Zhao, Zhou, Li et al. 2023. A Survey of Large Language Models. arXiv preprint.
Scope of the field we author strategy code with.
Kadan, Manela. 2026. The Value of Information: A Puzzle. arXiv preprint.
Information that does not price the way the theory says it should.
Imbalance3
Class imbalance, weighted losses, synthetic minority oversampling. Read because tradeable events are rare by construction, and a classifier that never fires scores extremely well on the metric nobody should be using.
2 entries listed · shelf of 3 files in Fig. 1
Chawla, Bowyer, Hall et al. 2011. SMOTE: Synthetic Minority Over-sampling Technique. arXiv preprint.
The reference oversampling method. Read alongside the fact that a classifier that never fires scores well on the wrong metric.
Boabang, Gyamerah. 2025. An Enhanced Focal Loss Function to Mitigate Class Imbalance in Auto Insurance Fraud Detection with Explainable AI. arXiv preprint.
Loss weighting instead of resampling.
Labeling2
Triple-barrier and event-based labelling. Two papers, because the no-leakage constraint decides most of this for us before the literature gets a say: any label whose definition reads data after the timestamp is out, whatever it is called.
3 entries listed · shelf of 2 files in Fig. 1
Kang. 2025. Stock Price Prediction Using Triple Barrier Labeling and Raw OHLCV Data: Evidence from Korean Markets. arXiv preprint.
Triple-barrier applied to raw bars on one named market.
Song, Liu, Chen. 2026. The Label Horizon Paradox: Rethinking Supervision Targets in Financial Forecasting. arXiv preprint.
The horizon choice as the decision that determines the result.
Kamat. 2026. RED-2400: A Public Benchmark of Algorithmically-Rejected Trading Events with Outcome Labels. arXiv preprint.
A benchmark of rejected events with outcomes attached. Rejections are the sample most research never keeps.
Cross-domain2
Cointegration and monitoring methods borrowed from structural and mechanical engineering. Kept as a standing reminder that the technique is rarely the novel part, and that a method's home field usually validated it more carefully than we will.
1 entry listed · shelf of 2 files in Fig. 1
Knes, Dao. 2024. Machine Learning and Cointegration for Wind Turbine Monitoring and Fault Detection: From a Comparative Study to a Combined Approach. Energies.
Cointegration used for fault detection. The technique is rarely the novel part, and its home field tested it harder.
On 2026-07-16 we ingested an external strategy corpus and scanned it with one question: does it hold a mechanism we are not already running? It does not. The scan is published because of what it found about our own ingestion, not because of what it found in the literature.
One ingested collection of 244 documents broke down roughly as follows. About a quarter was pure non-finance noise — immunology, DNA nanotechnology, weather models, EEG, high-entropy alloys, telescopes, RNA sequencing — pulled in by title resolution matching against reference lists. About a third was asset-pricing and market-structure theory rather than tradeable mechanism. Most of the remainder named mechanism families we already run. A separate compendium of 151 published strategies was, against the mandate as it stood that day, roughly four-fifths out of scope — the option-spread, fixed-income, credit, convertible, structured-product, real-estate and tax-arbitrage families — and every in-scope entry in it named a family already in our sweep corpus.
One candidate family came out genuinely unexplored. It is not named here. Naming it would say where we are about to look, and that is the one thing a reading list is not obliged to disclose.
A coincidence worth flagging so it is not read as an error: that external collection also contained 244 documents. It is a different collection, ingested separately for a different purpose, and it is not the corpus charted in Fig. 1.
The process failure and the fix are both published. The resolver matched titles against reference lists with no finance-relevance filter, so it imported whatever the citation graph handed it; a relevance filter was added at ingestion, and the instruction not to re-scan the collection expecting gold was written down with the reason. The same class of defect turned up in the metadata behind this page, and is the reason the lists are selections. Of 363 rows carrying a stored arXiv identifier, 325 pointed at a paper with a different title, so no stored identifier was used to build a link here: every arXiv link on this page was resolved against the arXiv API and then fetched. Stored DOIs held up — 201 of 201 verified, and stored author fields agreed with Crossref on 199 of 199. What is listed is the part of the corpus that survived that check, which is why it is 163 entries and not 802.
No entry on this page is marked as feeding a live book. Which shelf a live edge reads from narrows the search for anyone looking, so the association is not published: the shelf totals in Fig. 1 carry those papers, and no row on the shelves above is annotated in a way that would identify one. The two entry-level verdicts that do point somewhere — one Adopted, one Implemented and refuted — point at work already published in full on this site.
Two further omissions, stated for the same reason. Verdict assignment is a manual pass, complete on two shelves and not started on the rest, so eighteen shelves print citations and no verdicts; a hundred and forty-one rows of guessed verdicts would be worse than an honest blank. And the notes say what a paper settles for us without saying what it settles it at: where a number would be a calibration of our own screen, the shape is published and the value is withheld.
Each of those omissions is declared where it occurs rather than left to be noticed.
If we have misread a paper, tell us. Verdicts are the most correctable claim on this site — they are our judgement of our own record, and both halves of that can be wrong. Reports are acted on and logged in the library's corrections register like any other error.
Items here are never silently edited. A wrong item is corrected by a new item that cites it, so the record shows the correction as well as the correct answer.