Strategy grades & the assessment
Every backtest produces a grade and an assessment. This page publishes the complete formula for both, because a score you cannot recompute is a score you have to take on trust.
A, B+, B, C+, C, D. Earlier documentation described four (A through D). If you received a B+ or a C+ and found nothing explaining it, this is why.
The grade
A single letter from a 0-100 score summing five components.
| Component | Max points | What it measures |
|---|---|---|
| Return | 30 | Total return over the window |
| Sharpe | 25 | Risk-adjusted return |
| Drawdown | 25 | Worst peak-to-trough decline |
| Consistency | 15 | Profit factor and win rate together |
| Sample size | 5 | Number of trades |
Total: 100. A strategy with zero trades is graded D immediately, without scoring.
Grade thresholds
| Grade | Score |
|---|---|
| A | ≥ 85 |
| B+ | ≥ 72 |
| B | ≥ 62 |
| C+ | ≥ 52 |
| C | ≥ 42 |
| D | below 42 |
Return: 30 points
| Total return | Points |
|---|---|
| ≥ 20% | 30 |
| ≥ 12% | 24 |
| ≥ 5% | 16 |
| > 0% | 8 |
| ≤ 0% | 0 |
Sharpe: 25 points
| Sharpe ratio | Points |
|---|---|
| ≥ 1.5 | 25 |
| ≥ 1.0 | 18 |
| ≥ 0.5 | 10 |
| > 0 | 5 |
| ≤ 0 | 0 |
Drawdown: 25 points
| Max drawdown | Points |
|---|---|
| ≤ 8% | 25 |
| ≤ 12% | 18 |
| ≤ 15% | 12 |
| ≤ 20% | 8 |
| ≤ 30% | 4 |
| > 30% | 0 |
Consistency: 15 points
Profit factor and win rate must both clear a tier. This is why a high win rate alone does not earn the points, a 70% win rate with a profit factor of 0.9 is a strategy losing money slowly.
| Profit factor | AND win rate | Points |
|---|---|---|
| ≥ 1.5 | ≥ 60% | 15 |
| ≥ 1.2 | ≥ 50% | 12 |
| ≥ 1.0 | ≥ 45% | 8 |
| ≥ 0.8 | (any) | 4 |
| below 0.8 | 0 |
Sample size: 5 points
| Trades | Points |
|---|---|
| ≥ 100 | 5 |
| ≥ 40 | 4 |
| ≥ 20 | 2 |
| below 20 | 0 |
Only five points, but it is the component that stops a three-trade backtest scoring an A. With a perfect score on the other four a 12-trade strategy reaches 95, still an A, so read the trade count yourself as well. Five points cannot carry the whole burden of sample-size scepticism.
Grade A means the run scored 85 or more on the five components above, over the window tested.
It does not mean an edge was detected. It does not mean the strategy will work. It does not mean the platform recommends deploying it. It is a threshold score on past data: arithmetic, not a judgement. See Reading results honestly
The assessment
Alongside the grade, five descriptive labels. All are computed from the backtest; none is an assessment of you.
Return potential
| Label | Requires |
|---|---|
| Strong | Return ≥ 15% and Sharpe ≥ 1.0 and ≥ 20 trades |
| Moderate | Return ≥ 5% and ≥ 10 trades |
| Limited | Anything else |
The trade-count conditions are the interesting part: a 40% return on 6 trades is not "Strong", because six trades cannot support the claim.
Risk profile
| Label | Requires |
|---|---|
| Conservative | Drawdown ≤ 8%, volatility ≤ 12%, VaR(95) no worse than -2% |
| Moderate | Drawdown ≤ 15%, volatility ≤ 20%, VaR(95) no worse than -4% |
| Aggressive | Anything else |
A description of the strategy's behaviour, not of a suitable investor.
Drawdown tolerance required
| Label | Requires |
|---|---|
| Low | Drawdown ≤ 8% and recovered to the prior peak within 30 days |
| Medium | Drawdown ≤ 15% and recovered within 60 days |
| High | Anything else, including never recovering within the window |
Recovery time is why this is not just a restatement of drawdown. A 10% drawdown that recovered in three weeks and a 10% drawdown still unrecovered at the end of the window are very different experiences.
Market alignment
Whether the strategy's direction matched the market's behaviour over the window. A bull market is classified at ≥ +8% price return, a bear at ≤ -8%; the label is adjusted up at a win rate above 60% and down below 30%.
Useful for one specific question: did this strategy work, or did the market just go up?
Risk profile match
The label to read most carefully.
| Value | Computed when |
|---|---|
| Did not clear the return and Sharpe floor | Return ≤ 0 or Sharpe ≤ 0 |
| Suits a high drawdown tolerance | Aggressive risk profile, or High drawdown tolerance, or ≥ 20 trades/month |
| Suits a medium drawdown tolerance | Moderate profile, or Medium tolerance, or ≥ 8 trades/month |
| Suits a low drawdown tolerance | Everything else |
The engine produces this from the backtest's own numbers, and the UI has historically rendered it under the heading "Recommendation" with values naming investor categories, "Beginner Traders", "Experienced Traders", "Not Recommended".
That framing is wrong and is being changed. A label naming a category of person a strategy is "for" is a suitability assessment, and Stretus does not perform suitability assessments. It holds none of the information one would require.
What the label actually measures is what the strategy demands: its drawdown depth, its recovery time, its trade frequency. Read it as a statement about the strategy, never as a statement about you. Renaming the field and restating its values is a P0 item on the compliance roadmap.
Worked calculation
Illustrative. Not a real result.
A backtest returns: total return 14.2%, Sharpe 1.15, max drawdown 11.4%, profit factor 1.31, win rate 54%, 63 trades.
| Component | Value | Points |
|---|---|---|
| Return | 14.2%, clearing the 12% tier | 24 |
| Sharpe | 1.15, clearing the 1.0 tier | 18 |
| Drawdown | 11.4%, clearing the 12% tier | 18 |
| Consistency | PF 1.31 ≥ 1.2 and WR 54% ≥ 50% | 12 |
| Sample size | 63, clearing the 40 tier | 4 |
| Total | 76 |
A total of 76 gives Grade B+, which starts at 72.
Assessment: Return potential Moderate (14.2% is below the 15% Strong floor, by 0.8 percentage points). Risk profile Moderate. Drawdown tolerance depends on recovery time, which the components above do not show.
Note how close that Strong/Moderate boundary is. A tier boundary is a step function, and a strategy at 14.2% is not materially different from one at 15.1%. Read the numbers, not only the labels.
Using the grade sensibly
| Grade | A reasonable reading |
|---|---|
| A | Cleared every threshold comfortably on this window. Paper it |
| B+ / B | Cleared most thresholds. Identify the weakest component and make one change |
| C+ / C | Some components cleared, others did not. Read which before deciding whether to iterate |
| D | Either no trades, or it failed broadly. A different idea is usually better than patching this one |
Two strategies can both grade B and be entirely different problems. One with a great return and an ugly drawdown, one with a modest return and excellent risk control. The letter compresses that away. The component breakdown is on the backtest Overview tab.
Limitations
A grade describes one window. It is arithmetic on the tested period and nothing else.
Tier boundaries are step functions. 14.9% and 15.1% return are one point apart in reality and six in the score.
Sample size carries only five points. A small-sample A is possible. Read the trade count.
The grade cannot detect overfitting. A strategy tuned to the window grades beautifully on it. See Improving a strategy.
Costs are in the numbers, but assumptions are not. The grade uses cost-adjusted results, and those costs carry a model version. See Fees & charges
Next
- Performance metrics, what each figure means
- Reading results honestly
- Backtest limitations