Failures are data too.
Instead of collecting only attractive results, we document the mechanisms that can make a false result look convincing. These notes are based on issues and design principles encountered in real research, audits, and operations.
Even the same second can contain future information.
If second_ts, event_ts, and created_at are mixed, future events can leak into past features.
Why is 300 not the conclusion?
Multiple coins at the same timestamp, overlapping horizons, and shared market shocks can inflate nominal n.
UNKNOWN is not PASS.
Why unsupported areas remain UNKNOWN instead of being marked PASS.
Why not delete failed experiments?
Quarantine does not erase inconvenient data. It preserves the original record and incident history while excluding affected data from confirmatory analysis.
Why keep effect metrics blinded before 300?
Once you see the result, you also gain the freedom to adjust thresholds or cohorts in ways that favor it.
Why keep research records separate?
Why hypotheses, failures, changes, and evidence locations are also preserved outside the execution server.
Why do fees and slippage inflate short-term backtests?
Why seemingly small fees and slippage can materially change the net result of a high-turnover strategy.
How can an 80% win rate still lose money?
Why a high win rate alone says little about profitability without average win, average loss, and expectancy.
Why maximum drawdown (MDD) comes before headline return.
Why identical final returns can hide very different loss paths and operational risk.
Why separate in-sample and out-of-sample?
Why using the same data to design and evaluate a strategy can overstate performance, and how a clean out-of-sample split helps prevent that.
Why backtest fills differ from real fills.
Why visible chart prices can diverge from executable prices once order-book depth, queue position, latency, and partial fills matter.
Why repeated backtest tuning leads to overfitting.
Why repeated tuning to historical data can improve in-sample performance while weakening performance on new periods.
Should the stop be -1% or -5%? The trap of choosing a single fixed number.
This explains why the same -5% can be a different risk if the fixed stop loss is not linked to volatility, trading costs, or position size.
More trades do not just create more opportunities; they also create more costs.
How fees, spreads, slippage, and execution uncertainty compound as trading frequency rises.
Why is walk-forward validation harder than a single OOS split?
How walk-forward validation repeatedly advances the training and test windows, exposing regime changes and the consequences of re-optimization rules.
Why does the past look better when we test only coins that survived?
Why survivorship bias appears when a crypto backtest includes only coins that are still listed today and excludes those that disappeared.
Why an AI-discovered strategy can still be rejected immediately.
Why check point-in-time, leakage, and reproducibility before high returns.
Why a popular crypto strategy should not be trusted at face value.
Even famous strategies must have their rules fixed and re-verified with OOS, transaction costs, and execution reality.