Methodology
How we score strategies
The eight checks every strategy on The Algo Bench is scored against, how each one is decided, and how we rate risk.
Every strategy carries three tags at the top: whether it is investing or trading, how many of our eight checks it passed, and its risk level. There is no star rating and no one-word verdict: the checks say what held up, the risk tag says what it would have felt like, and the write-up explains the rest.
These checks describe how a strategy behaved on historical data. Passing all eight does not mean a strategy will make money in future, and none of this is a recommendation to trade anything. See the disclaimer.
Investing or trading
- Investing: stock or ETF positions held for weeks to years, bought outright (delivery), without leverage.
- Trading: intraday, overnight or short swing trades, and anything in futures or options.
The eight checks
Each check is marked ✓ passed, ✗ failed or – not tested. Not tested counts as not passed: we would rather show a gap than guess.
1. Profitable after costs, beyond luck
Makes money after brokerage, taxes, charges and slippage, by more than luck could explain.
Passes when: Two things, both required. The net result is positive at our standard cost model. And it is not plausibly luck: where it can be measured, the t-statistic of the average net trade is about 2 or higher, meaning there is only about a 1-in-20 chance that a strategy with no real edge would look this good.
Why it matters: Most intraday ideas have a real but tiny gross edge that costs wipe out.
2. Beats a simple alternative
Does better than doing something much simpler.
Passes when: Investing strategies: over the same period it beats simply buying and holding the index (the NIFTY 500 for stock strategies). Trading strategies: its return on capital (see below) beats leaving the same capital in a low-risk deposit, taken as 7% a year (a bank fixed deposit or liquid fund).
Why it matters: A positive return means little if a much simpler or safer choice did better: holding the index for an investor, or a deposit for the capital a trader has to set aside.
3. Holds up out of sample
Still works on data it was never tuned on.
Passes when: It stays profitable on a period that played no part in designing or tuning the rules: a held-back period, a later period, or a walk-forward test.
Why it matters: Rules tuned on one stretch of history often fit its noise, not a lasting effect.
4. Not fragile to small parameter changes
Nearby settings work too: a plateau, not a spike.
Passes when: Settings close to the chosen ones (a slightly different lookback, threshold, stop or timeframe) are also profitable, by the same standard as check 1. Results form a broad plateau rather than a single lucky peak. This check can only pass if check 1 passed.
Why it matters: If only one exact setting works, the strategy has probably been fitted to the past.
5. Consistent over time
Profitable in most quarters (or years), not one good stretch.
Passes when: Trading strategies: profitable in at least 65% of calendar quarters, or at least 85% for option-selling strategies, which should be steadier. Investing strategies: profitable in at least 65% of calendar years. Every quarter (or year) in the test counts, including ones spent out of the market. The worst quarter (or year) must not be out of proportion to the typical gain.
Why it matters: A good total can hide a curve carried by one or two outlier periods.
6. Survives higher costs
Still profitable if trading costs turn out worse.
Passes when: It remains profitable at double our standard cost assumption (or at the worst realistic cost for the instrument, if that is higher). This check can only pass if check 1 passed: there has to be a profit to survive.
Why it matters: Real fills are often worse than assumed, especially for frequent trading; thin edges die first.
7. Profit not concentrated in a few trades
The profit doesn't depend on a handful of lucky trades or periods.
Passes when: It is still profitable after removing its 5 best trades, and no single quarter (year, for investing strategies) accounts for more than a third of the total profit.
Why it matters: Profit carried by a few outliers is a classic sign of overfitting, and is unlikely to repeat.
8. Enough trades and years of data
Tested on enough history to mean something.
Passes when: Trading strategies: at least 100 trades over at least 3 years. Investing strategies: at least 10 years of data, including at least one bear market.
Why it matters: A handful of trades or a single market phase can't separate skill from luck.
Before any check: the house rules
A result is only scored at all if it was tested honestly: on an instrument you can actually trade, with realistic fills, without look-ahead, and on survivorship-free data. A result that fails these isn't a strategy result; it's a Landmine. Details on the Methodology page.
Capital and return (trading strategies)
Points alone don't tell you much, so trading strategies are also shown per lot of 65 units, in rupees:
- Capital = the exchange margin for one lot plus the strategy's worst backtest drawdown for one lot, the buffer you would have needed to survive its worst stretch without topping up. For option-selling strategies we use the worst intraday loss instead, because margin calls happen during the day.
- Return on capital = average annual profit per lot ÷ that capital. For investing strategies the equivalent is CAGR.
- Worst drawdown as a % of capital, so you can see how much of the account the worst stretch consumed.
The current lot size is applied to the whole history, although it has changed over the years. Margins are approximate and change with volatility and exchange rules.
Risk
Risk is not one of the eight checks. People tolerate losses very differently, and the same strategy can be reckless with most of your capital yet reasonable with a small slice of it. So risk doesn't decide whether a strategy is good or bad. Instead, every strategy shows its worst drawdown (the largest fall from a peak) and a risk level based on worst drawdown ÷ average annual return: roughly, how many years of average returns that worst fall would take to recover. This works the same for investing results in % and trading results in points.
| Risk level | Meaning |
|---|---|
| Low | The worst drawdown was less than six months of average returns. |
| Medium | The worst drawdown was six months to a year of average returns. |
| High | The worst drawdown was one to two years of average returns. |
| Very high | The worst drawdown was more than two years of average returns. |
| Extreme | Losses are not capped. A single extreme day can exceed anything seen in testing, so the historical drawdown understates the risk. |
| Not rated | The strategy loses money after costs, so there is no return to measure risk against. |
Strategies whose losses are not capped (for example, selling options without protection) are always rated Extreme, whatever their backtest drawdown: a few years of history can't show the worst day such a strategy can have.
Where relevant, a strategy also carries risk flags:
- Always fully invested: carries full market risk, with no cash cushion.
- Profit depends heavily on a few large trades or periods.
- Uses leverage or margin: losses can exceed the capital at risk.
- Hard to trade at larger size without moving prices.
- No stop-loss: a position can keep losing until the exit rule triggers.
- Holds positions overnight: exposed to gaps that no stop-loss can prevent.
- Short options without protection: losses are not capped.
The risk level describes the strategy, not who it suits. Whether a level of risk is acceptable depends on your own circumstances.