Lesson 6: How to Evaluate a Trading Strategy
- 6 days ago
- 6 min read
Looking Beyond Profit to Understand Risk, Consistency and Robustness
A profitable backtest is only the beginning of strategy evaluation. The more important question is how that result was produced: how much risk was taken, how consistent the outcomes were, whether the result depends on a small number of trades, and whether the behaviour remains credible when assumptions or market conditions change.
This lesson introduces a practical framework for reading strategy statistics together rather than treating one headline number as proof of quality.

The objective is not to find the strategy with the highest historical return. It is to understand the quality and reliability of the return process.
1. Start With the Return — But Do Not Stop There
Net return describes the overall historical outcome after the costs included in the test. It is useful, but it says little about the path taken to reach that result.
A strategy that gains 20% with a relatively controlled path is not equivalent to one that gains 20% after experiencing severe losses, unstable exposure or one unusually successful trade.

2. Maximum Drawdown
Drawdown measures the decline from a previous equity peak to a later trough. Maximum drawdown is the largest such decline during the evaluation period.

Drawdown helps answer a practical question: how much deterioration did the strategy experience while producing its return? It should be interpreted alongside leverage, position sizing and the length of the recovery period.
Recovery Matters Too
Two strategies with the same maximum drawdown may still feel very different if one recovers quickly while the other remains below its previous peak for a long period. Duration and depth are therefore both useful when reviewing downside behaviour.
3. Win Rate: Useful but Easy to Misread
Win rate is the percentage of completed trades that generated a positive result. A high win rate can look attractive, but it does not show the size of winning trades relative to losing trades.

A strategy can win frequently and still lose money if occasional losses are large. Conversely, a strategy can lose more often than it wins and remain viable if its winning trades are sufficiently larger than its losing trades.
4. Average Win, Average Loss and Payoff Ratio
Average win and average loss provide the missing context behind win rate. Their relationship is often described through a payoff or reward-to-risk ratio.
For example, a strategy with an average winning trade of 2R and an average losing trade of 1R has a 2:1 average payoff relationship. This does not guarantee profitability, because the frequency of wins and losses still matters.
5. Expectancy: Combining Frequency and Payoff
Expectancy estimates the average historical outcome per trade by combining the probability and size of winning and losing outcomes.
Expectancy = (Win Rate × Average Win) − (Loss Rate × Average Loss)
Suppose a strategy wins 45% of trades, earns an average of 2R on winners and loses 1R on losing trades. Its illustrative expectancy is:
0.45 × 2 − 0.55 × 1 = +0.35R per trade
Expectancy is useful because it connects win rate and payoff into one framework, but it remains a historical estimate and depends on the quality and size of the sample.
6. Profit Factor
Profit factor compares total gross profits with total gross losses.

A profit factor above 1 means gross historical profits exceeded gross historical losses under the assumptions of the test. The further it is above 1, the larger the historical margin between the two. However, a very high value from only a small number of trades should be treated cautiously.
7. Sharpe Ratio and Risk-Adjusted Performance
The Sharpe ratio is commonly used to compare return with the variability of returns. In simplified terms, it asks how much excess return was generated relative to the amount of volatility in the return series.
It can be useful when comparing strategies with different return paths, but it is not a complete measure of risk. It relies on assumptions about the return series and does not describe all forms of downside risk, tail events or drawdown behaviour.
Risk-adjusted metrics should complement drawdown analysis, not replace it.
8. Trade Count and Sample Size
Performance statistics become harder to interpret when they are based on very few observations. A strategy with ten trades can show an impressive win rate or profit factor simply because the sample is small.
Trade count should therefore be considered whenever reviewing a metric. The relevant question is not only “What is the number?” but also “How much evidence produced that number?”
9. Read Metrics Together

No individual statistic captures the full behaviour of a trading system. Win rate describes frequency, average win and loss describe payoff, expectancy combines the two, drawdown describes downside experience, and risk-adjusted measures describe another dimension of consistency.
The strongest evaluation comes from looking for a coherent story across the metrics rather than searching for one impressive figure.
10. Exposure and Capital Efficiency
A strategy’s return should also be viewed in relation to how much capital and time were exposed. A strategy that is invested only occasionally may have different portfolio value from one that continuously consumes the same capital.
Exposure can be measured in different ways, including percentage of time in the market, average gross exposure, maximum simultaneous positions or capital allocated to open trades.
11. Consistency Across Time
A strategy that earns nearly all of its historical profit in one short period may be less convincing than a strategy whose edge appears across several independent periods.
Useful checks include reviewing yearly or monthly performance, rolling statistics, different volatility environments and whether the strategy depends on one unusually favourable market episode.
12. Robustness Across Market Regimes
Strategies are rarely equally effective in every environment. Trend-following logic may behave differently in range-bound markets; mean-reversion systems may struggle when strong directional moves persist.

The goal of regime analysis is not to demand positive performance everywhere. It is to understand the conditions under which the strategy’s assumptions are most and least reliable.
13. Parameter Stability
A strategy should also be checked around the selected parameter values. If performance disappears after a very small parameter change, the historical result may be fragile.
Stable regions are generally more informative than a single historical optimum. This is one reason parameter sensitivity testing is an important part of strategy evaluation.
14. Evaluate After Realistic Costs
All key metrics should be reviewed after reasonable execution assumptions have been included. A strategy can have attractive gross expectancy but little usable margin once spread, commission, slippage and financing costs are considered.
Cost sensitivity is particularly important for high-turnover strategies or systems targeting relatively small average moves.
15. Out-of-Sample Performance
A strategy that performs well only on the data used to build it provides limited evidence of generalisation. Out-of-sample testing evaluates the completed strategy on historical data that did not influence the original design decisions.
The unseen result does not need to match the development period perfectly. More important is whether the strategy retains broadly credible behaviour rather than collapsing once it leaves the data used to construct it.
16. A Practical Strategy Scorecard

A scorecard helps prevent selective attention. Instead of asking only whether a strategy made money, review whether the result survives realistic costs, whether drawdown is compatible with the intended risk level, whether expectancy is supported by enough observations, and whether the behaviour remains stable outside one narrow historical configuration.
17. Questions to Ask Before Accepting a Backtest
Is the return calculated after realistic trading costs?
What was the maximum drawdown and how long did recovery take?
How many trades produced the statistics?
Are average wins and losses consistent with the reported win rate?
Is historical expectancy positive and reasonably stable?
Does one trade or one period explain a large share of the total result?
How does performance change across different market regimes?
Do nearby parameter values produce broadly similar behaviour?
Does the strategy remain credible on genuinely unseen data?
Would the intended position sizing make the observed drawdown acceptable in practice?
Key Takeaways
Historical profit alone is not enough to evaluate a strategy.
Maximum drawdown describes an important part of the risk path.
Win rate must be interpreted together with average win and average loss.
Expectancy combines trade frequency and payoff into an average historical edge.
Profit factor compares gross profits with gross losses.
Sharpe ratio provides one view of risk-adjusted performance but does not replace downside analysis.
Trade count matters because small samples can produce unstable statistics.
Robustness should be checked across time, market regimes, parameters and unseen data.
Strategy evaluation is strongest when multiple metrics tell a consistent story.
Educational Notice: This material is provided for educational purposes only. The statistics, examples and charts are simplified illustrations of quantitative strategy evaluation and do not constitute investment advice, a personal recommendation or a guarantee of performance. Historical, simulated and backtested results do not guarantee future outcomes. Trading, particularly with leveraged products, involves significant risk.




Comments