Quant research infrastructure & backtest audits
quantlab: a futures research framework
A futures research and backtesting framework on Nautilus Trader, built on the methods of López de Prado's Advances in Financial Machine Learning. The goal isn't "find a profitable strategy" — it's making every conclusion hold up.
- 16 yearsCME minute data: 2,211 contracts, 164M bars
- 2,135Automated tests; ~29k lines of library code
- 81.75 GB → 0.10 GBPeak memory of sample-uniqueness weighting
- 0Bar-by-bar mismatches vs. the Nautilus engine (controlled test data)
- CME minute data
- Contract cleaning & continuous series
- Information-driven bars
- Features & event sampling
- Triple-barrier labels
- Cross-validation & backtests
- Conclusions under pre-registered criteria
Background
My early CTA strategies came from parameter optimization, and they fell apart out of sample. So I switched to the validation methods of Advances in Financial Machine Learning and built this framework, with tests guarding every layer: data, labels, cross-validation, backtests.
What I did
- Ingested 16 years of CME minute data; built continuous futures (the ETF trick), information-driven bars, features and event sampling.
- Implemented purged K-fold, combinatorial cross-validation, walk-forward and stress tests, plus triple-barrier labels and meta-labeling.
- Ran research with pre-registered criteria, provenance checks, multiple-testing corrections and independent reviews.
Problems and solutions
- Three traps in long histories: single-digit-year contract codes repeat every decade; codes get reassigned to a new contract in their expiry month; the same
instrument_idis reused for other products across decades (silver data would land in S&P futures). Contracts are now identified by code plus validity interval, with a guard before anything is written. - Bit-identical continuous futures: 1,386 rolls across 16 roots, identical whether built offline or driven live bar by bar.
- Batch/live parity: 113 features match bit for bit between batch and online computation; research-side backtests match the Nautilus engine bar by bar on controlled test data.
- Memory: replaced a dense matrix with a difference array and prefix sums for sample-uniqueness weights — from 81.75 GB to 0.10 GB per process.
- Upstream bug: found that Nautilus Trader 1.222’s catalog query silently ignored its bar-type filter.
The honest result
Across 12 CME futures and 16 years, I tested 11 signal families (77 configurations) and 6 trend filters under pre-registered criteria. None showed a robust, significant excess return over a passive long benchmark, so nothing went live. That is what a backtest audit is for: catching overfitting and leakage before real money is at risk.
What I can do for you
- Data pipelines for futures or equities: ingestion, cleaning, continuous contracts, validity checks
- Backtest audits: look-ahead, overlapping samples, cost assumptions, multiple testing
- Backtests and strategy setup on Nautilus Trader
Stack
Python · pandas · NumPy · Numba · scikit-learn · PyArrow / Parquet · Nautilus Trader · Databento · pytest