Skip to content
QuantWright

Reading results

A result shows what the strategy did on past data, how far that evidence can be trusted, and anything that does not add up.


The result

Every backtest produces a result card: the headline figures, the equity curve, the Honesty Score and a one-line verdict. The notes under it say what was measured, what was assumed, and what was not modelled. Ask about any figure in the chat and the answer is read from the stored run.

On a short test window the annualised figures (annual return, the annualised Sharpe and Sortino, Calmar, trades per year) are withheld rather than extrapolated from a span too short to annualise.

The tested period

The line under a result’s headline names the market, the bar size and the dates it tested, for example SPY · daily · 11 Oct 2018 – 28 Sep 2026 · 7.9 years. When the test is shorter than the depth or the dates you asked for, the reason is written beside the dates: the history starts later, or one run cannot hold that many bars.

Under the headline figures, each calendar year of the test shows its return and the trades entered that year. A first or last year the test covers only part of is labelled with the date it starts or stops. The returns are the Year column of the monthly returns in the Backtest tab.

The Period buttons on a result re-run the same code over the last 1, 2 or 5 years, Max, or two dates you pick. Max reaches up to 10 years on intraday bars, less where the data starts later or one run's limit comes first, and all daily history on daily bars. Each button says how far it reaches before you press it. They send /period to the chat. Nothing is regenerated: the code, its parameter values and its costs stay the same, and the new card says which version’s code it ran.

Once the same strategy has run over more than one period, its newest result lists them side by side under This strategy across periods: each window’s net return, maximum drawdown, trades, win rate and Honesty grade, with a link to its card. Only runs of the same code, parameter values and costs are listed; a version the AI rewrote is a different strategy and is left out.

Returns and risk

Total return

Cumulative profit over the whole backtest, after costs, as a percent of starting capital.

Annual return

Return compounded to a per-year rate, so periods of different length compare fairly.

Benchmark

Buy-and-hold of the same asset over the same window, from its first close to its last, as a price return (dividends not included, as on the strategy side) — the bar your strategy has to beat to be worth the effort.

Time in market

The share of the tested bars on which the strategy held a position. A strategy that is out of the market most of the time is not exposed to it, which matters when comparing it with buy-and-hold.

Max drawdown

The largest peak-to-trough drop in account value — the worst losing streak you'd have had to sit through.

Longest drawdown

The longest stretch, in calendar time, from an equity peak until the account was back at it, on the same account curve as the max drawdown. One still open at the end of the test runs to the last bar.

Average drawdown

The mean depth of the drawdowns on the account curve, each measured from its peak to its lowest point.

Average drawdown length

The mean length, in calendar time, of the drawdowns on the account curve, peak to recovery.

Sharpe ratio

Per-trade Sharpe, annualised: the mean of the per-trade net returns (on the traded price, after costs) divided by their standard deviation, times √(trades per year). Risk-free rate 0. It is not the textbook Sharpe on daily returns, so the usual rules of thumb for that one do not carry over: a strategy that trades often can print a large figure from a small edge per trade.

Sortino ratio

Like Sharpe, on the same per-trade basis (× √(trades per year), risk-free rate 0), but only downside deviation below zero counts against you — upside swings don't. Higher is better.

Calmar ratio

Annual return divided by worst drawdown. Higher means more reward per unit of pain; above ~1 is healthy.

Equity curve

Your account value over time. Smooth and rising is good; a flat stretch then one spike is a warning.

Monthly returns

Each cell is that calendar month's return. Scan it for consistency — a few huge months carrying everything is fragile.

Trade statistics

Trades

Number of completed round-trip trades. Fewer than ~30 and the statistics aren't reliable.

Win rate

Share of trades that made money. A high win rate with tiny wins and huge losses can still lose — read it alongside profit factor.

Profit factor (price)

The winning trades’ returns added up, divided by the losing trades’ returns added up — each trade’s return measured on its own traded price, after costs, not in dollars on the account. It matches the dollar ratio only when every position is the same size; above 1, the winners outweighed the losers trade for trade.

Expectancy per trade

Average return per trade, measured on each trade’s traded price after costs. Positive means the trades made money on average over this test, which by itself does not show a real edge.

Average win

Mean return on the trades that won.

Average loss

Mean return on the trades that lost (shown negative). Winners should comfortably outweigh losers.

Largest win

The best single trade’s return, measured on its traded price after costs — the same basis as expectancy and the average win.

Largest loss

The worst single trade’s return, measured on its traded price after costs — the same basis as expectancy and the average loss.

Payoff ratio

Average win divided by the size of the average loss, both per-trade returns on the traded price. Read it with the win rate: a low win rate needs a high payoff ratio.

SQN

System Quality Number: √(number of trades) × the mean per-trade return ÷ its standard deviation, on the per-trade price returns after costs. Van Tharp’s original uses R-multiples, so the two are not the same scale.

Long / short

The trades split by side: how many, the share that made money, and the mean per-trade return on the traded price after costs.

Max consecutive wins

Longest winning streak — useful for keeping expectations grounded during a hot run.

Max consecutive losses

Longest losing streak. The real gut-check: could you keep trading the plan through it?

Trades per year

How often the strategy trades, annualized. This drives how much commission and slippage you actually pay.

Holding period

How many bars each trade stayed open — tells you whether this is scalping, swing, or position trading.

Trade distribution

Histogram of per-trade outcomes. A tight cluster is dependable; a few outliers carrying the result is fragile.

The robustness check (/validate)

Robustness check (/validate)

The run’s own trades re-split many ways into training and test folds across its history (CPCV), with a gap between them so nothing leaks across. Nothing is re-fitted on the training folds and every trade lands in some test fold, so this shows how steady the fixed strategy is across stretches of its history. It is not walk-forward optimisation and not a test on data the strategy has never seen.

CPCV (cross-validation)

Combinatorial Purged Cross-Validation: splits the history into many train/test fold combinations with a gap between them so future data can’t leak into the past. Here it re-splits one fixed strategy’s trades and re-fits nothing, so it shows how steady they are across the history; it is not evidence the edge isn’t curve-fit.

Test-fold Sharpe (per trade)

Per-trade Sharpe over the trades in the test folds. A fixed strategy is not re-fitted on the training folds and every trade lands in some test fold, so this equals the full run’s per-trade Sharpe. It is not measured on data the strategy has never seen.

Training-fold Sharpe

Sharpe over the trades in the training folds. Nothing is fitted on them: they are the same strategy’s trades from other stretches of the history.

Degradation

How the Sharpe changes from the training folds to the test folds. Nothing was re-fitted, so a drop is variation across the history, not by itself proof of overfitting.

Deflated Sharpe Ratio

Shown 0–100%: the probability the strategy’s true Sharpe beats what the best of the variations tried would show by luck. Testing many ideas inflates luck; this deflates it back out. It is computed on this run’s own trades, not on unseen data. Higher is better.

Probabilistic Sharpe (PSR)

The probability your true Sharpe is above zero, given how many trades you have and how lumpy they are. Climbs toward 95% as more profitable trades accumulate.

Windows

Number of separate train/test splits of the run’s own history the check used.

Test-fold trades

Trade count in the check’s test folds. For one fixed strategy that is every trade of the run. Too few and the result is just noise.

Overfitting

Tuning a strategy so tightly to past data that it memorizes noise instead of a real edge — it looks great in the backtest and falls apart live.

Monte Carlo

Monte Carlo

We redraw your trades at random — with replacement, in blocks of consecutive trades, so some repeat and some are left out — thousands of times, to see the spread of returns and drawdowns those same trades could give. It takes the trades as they are, so it does not test whether the edge is real. Use "Run Monte Carlo" under the result for the deep 10k-path version.

Profitable paths

Share of simulated runs that ended in profit — the run’s own trades resampled WITH REPLACEMENT, so each path varies in which trades it draws. Which trades a path draws decides where it ends; the order they come in does not (it only changes how deep the path dips on the way). It assumes the edge is real: a curve-fit strategy scores high here too, so this figure is no evidence that the edge exists. Not the chance of making money in future. The Honesty score says how much evidence there is; /validate is not an out-of-sample test.

Median return

The middle outcome across all simulated paths — half do better, half worse. More honest than the single backtest figure.

Terminal return

Final return at the end of each simulated path. The p5/p95 range shows a realistic best-to-worst spread.

CAGR

Compound annual growth rate — each simulated path's return annualized over the CALENDAR span of the backtest window, the same basis buy-and-hold is quoted on. The p5/median/p95 spread shows the realistic range of yearly compounding, not a single lucky number.

CVaR (Expected Shortfall)

The AVERAGE outcome across the worst 5% of simulated paths — not just the p5 boundary. It answers "when it goes bad, how bad on average?" and is the metric risk desks size to.

Median max drawdown

Across all simulated paths, the typical worst drop. Half of outcomes are deeper than this, half shallower.

Worst max drawdown

The deepest drop seen across every simulated path — a stress test of the ugliest realistic run.

Risk of ruin

Share of simulated runs whose worst drawdown reached a severe threshold (the deepest measured is ~50%) — the run’s own trades resampled WITH REPLACEMENT, so each path varies in which trades it draws and in the order they come, which matters here because a drawdown depends on the sequence. It assumes this run's edge is real and repeats, so it is not the probability of wiping out your account in future.

The Honesty Score

A 0 to 100 score of how much the evidence supports the result. It marks down what usually makes a backtest look better than it is: too few trades, a Sharpe ratio that could be luck, a search over many versions, and returns that are mostly the market. It is not a forecast.

ComponentWeightWhat it measures
Sharpe confidence40%The probability that the strategy’s true Sharpe ratio is above zero, given how many trades there are and how lumpy they are (Probabilistic Sharpe).
Sample size15%The number of trades behind the result; after /validate, the trade observations across its test folds. Full marks at 100.
Track record length10%Trades on record against the minimum needed for this Sharpe ratio to be statistically significant (MinTRL).
Market independence10%How much of the return the market explains: the R² of the strategy against SPY. Measured by /validate. Not scored for a strategy that trades SPY itself.

When you try several versions of a strategy in one chat, the Sharpe check is deflated for the search: its 40% splits into Deflated Sharpe (30%) and Probabilistic Sharpe (10%), and the score’s breakdown says which variant this is and what the best of that many random strategies would show by luck.

The remaining 25% belongs to fold stability, an overfitting check that needs competing parameter sets and is not computed yet. It is left out rather than counted as a pass, and its points are not handed to the other checks. So the checks allow at most 75, and 65 for a strategy that trades SPY itself, so no run can reach grade A (80 or more) until fold stability is measured. On top of that, every score is capped at 64 (grade C) until a held-out test has run; see caps for how to run one.

Grades

GradeScore
A80 to 100 (not reachable today)
B65 to 79
C50 to 64
D35 to 49
F0 to 34

Caps

Some findings do not lower the score a little; they put a ceiling on it, because the number would otherwise answer a different question. The verdict names the cap whenever one applies.

64 (grade C)

Capped at grade C until a held-out test has run: a test on data the strategy was not developed or chosen on. To run one, put a hold-out on the record before you develop (“keep everything from 2024-10-01 out of sight”), test only on earlier years, then re-run the card you choose, unchanged, over exactly those years with /period (for example /period 2024-10-01 2026-10-01 v5, the card’s period control, or “re-run v5 unchanged on the held-out years”). That run is the test, scored on its own trades: it lifts this cap only if no earlier run on that market in the chat had tested any of those years, and only once. It is not deflated by the runs it was picked after, since none of them saw those years; only by held-out attempts. A second try, a changed version, or any later run over those years does not lift it, and like every run it counts as a trial. A run that took no trades because a date in its code shut those years showed nothing, so it does not use them up. Where a check is still unmeasured on the held-out run (market independence), /validate on that card measures it. /validate on any other run is not a held-out test: it re-splits the run’s own trades into folds, re-fits nothing, and every trade lands in a test fold, so it measures how steady the result is across its history and does not lift this cap.

49 (grade D)

The edge cannot be told apart from zero or from luck: the bootstrap 95% confidence interval for the Sharpe ratio includes zero, or, after several versions were tried, the deflated Sharpe is 80% or lower.

64 (grade C)

The backtest departs from the strategy you described in part: for example, a session window you set that some trades fall outside.

39 (grade D)

The backtest does not implement the strategy you described: for example, a rule never to hold overnight, on a run where every trade was held overnight.

Verdict words

The headline on each result is one of these:

Pass + Edge

Every check has run on a held-out test (see caps), and the score is 65 or more.

Weak Edge

Every check has run, and the score is from 35 to 64.

Edge unverified

A check is still to run (send /validate to measure the rest), and the score so far is 35 or more.

Insufficient evidence

No trades, or a score below 35.

Account wiped out

The account lost everything during the test. This is checked before anything else, since a ruined account ends the test early and leaves few trades.

Warnings

Warnings sit with the result they are about. Heads up notes give context: a default that was used, a data caveat, what a figure is based on. Warning notes flag something that changes how the result should be read.

Routine notes fold into one row that says how many there are. The loudest ones, such as a rule that never fired or a result that contradicts its own code, never fold. Some warnings carry a button that applies the fix and runs again.

Look-ahead is code that reads bars it could not have seen yet; its results cannot be achieved. Every run’s code is checked against a list of known look-ahead patterns. A held-out test is also re-run with the bars after 8 cut points (20% to 70% of its window) removed, and so is every run of code you supplied or changed (your own file, /runcode, a Code-tab edit, a Pine translation) after 4 cut points (27% to 70% of its window). Code that reads only the past cannot change a trade it took before a cut, so if any such trade changes, a warning names the first by date and bar, and the run is marked invalid and not scored. A held-out test whose years closed fewer than 10 trades before the latest cut is checked over the hold-out together with the history fetched before it: look-ahead is a property of the code, not of the years. The grade-C cap stays when this check cannot run: still too few trades, a run made in segments, a shortened re-run that failed, or code whose identical runs take different trades. The check cannot see a leak that changes no trade before a cut: a high of the whole series reached before the first cut is the same high on every shortened run, and a leak that reads only a few bars ahead shows only on trades next to a cut, so it can fall between them.

More on the checks behind the numbers: Market data and FAQ.

Not answered here? Write to us through the Support page.