Once a strategy has rules and out-of-sample results, you have to judge it — and judging only by total return is how people fall in love with fragile, dangerous strategies. A proper evaluation weighs return against risk and consistency, and stays alert to the ways you can fool yourself.
The metrics that matter
- Expectancy — average profit per trade (in R); is it positive net of costs? The first gate.
- Maximum drawdown — the worst peak-to-trough fall; can you survive and stomach it?
- Risk-adjusted return (e.g. Sharpe) — return per unit of volatility; rewards smooth gains over wild ones.
- Sample size — enough trades for the result to mean something (a 5-trade 'edge' is noise).
- Consistency — does it work across regimes, or only in one lucky environment?
Return is not the whole story
A strategy returning 40% with an 60% drawdown is usually worse than one returning 20% with a 10% drawdown — the second is survivable and repeatable; the first will likely ruin you before it compounds. This is why risk-adjusted metrics and drawdown matter more than the headline number, and why this project reports Sharpe and max-drawdown beside returns. A high return born of huge risk is not skill; it is a bet that hasn't lost yet.
The honesty test
Finally, interrogate your own result. Is the sample big enough? Did you optimise until it looked good (overfit)? Did you include all costs? Would it have survived a different period? The goal is not to find a reason to believe — it is to find the reasons it might be wrong, and only trust what survives them. That adversarial honesty is precisely the discipline behind this project's published 'no demonstrated edge' conclusion: they tried hard to break their own strategies, and reported what they found. Evaluate yours the same way.