Skip to content
← All financial terms

Planning & advice · Financial term

Historical Backtesting

Also called backtesting · backtest · historical simulation · rolling-period analysis · backtested performance

What is historical backtesting?

Historical backtesting is testing a strategy on real past market data as if it had been followed at the time. In retirement planning, it usually means running a withdrawal plan through every stretch of history of the same length, such as each 30-year period since 1926, and counting how many periods the money lasted. The result shows how a strategy would have fared, not how it will.

8 min readWorked example4 common questions

How a rolling-period backtest works

The standard retirement backtest uses rolling periods. You pick a strategy, such as withdrawing 4% of the starting portfolio and raising the dollar amount with inflation each year, and a length, such as 30 years. Then you run the strategy starting in the first year of the data, again starting one year later, and so on until the last start year that still has a full 30 years of data behind it.

William Bengen’s 1994 paper used this method on US data from 1926, starting a hypothetical retiree on January 1 of each year. The Trinity study of 1998 did the same with 1926–1995 data, using the S&P 500 for stocks and long-term high-grade corporate bonds.

Each period keeps history in its original order. Returns, Inflation and the way stocks and bonds moved together all arrive exactly as they did, through the Great Depression, the high-inflation 1970s and the recoveries after them. That is the method’s great strength: the market path itself is not an assumption.

How to read a backtest and fit it to your plan

A backtest produces one outcome per starting year, and the way you summarize them changes the story. Bengen reported the worst case: with a rebalanced half-stock, half-Treasury portfolio, a 4% first-year withdrawal raised each year for inflation always lasted at least 33 years, while 4.25% could run out in 28. That became the 4% rule. Yet in 40 of the 51 start years he charted, 4% lasted the full 50 years his charts showed. A few bad periods set the rule, chiefly retirements that began in the late 1960s, just ahead of the 1973–1974 bear market and high inflation, a textbook case of sequence of returns risk.

The Trinity study reported a success rate instead: the share of periods that ended with money left. A success rate hides how badly the failures failed and how close the survivors came, while a worst case rests on a single period, so read both. Studies of the safe withdrawal rate still quote results both ways.

Both studies tested a generic retiree with rigid, inflation-adjusted withdrawals and judged success mainly at 30 years. A backtest tells you more when it matches your own plan:

  • Use your own horizon. A 40-year test fits only 31 rolling periods into 70 years of data, 10 fewer than a 30-year test, and longevity risk argues for running past your life expectancy.
  • Use your real income timing, such as a pension or a Social Security start at 70, instead of one flat withdrawal.
  • Apply the spending rule you would really follow, such as dynamic spending cuts after bad years.
  • Record each period’s lowest balance and ending balance, not just whether it passed.

Why backtests look sturdier than they are

A backtest is only as broad as the history behind it, and that history is shorter than it looks. Rolling periods overlap heavily: the 30-year periods starting in 1926 and 1927 share 29 of their 30 years. Seventy years of data yield 41 rolling 30-year periods but only two that do not overlap at all, so a 95% success rate rests on far fewer independent outcomes than the count suggests.

The record is also one country’s, and the most recent start years cannot be judged until their full period has passed: a retirement that began in 2000 will not finish a 30-year test until the end of 2029.

Above all, the future is under no obligation to stay inside the range of the past. FINRA’s communications rule bars broker-dealers from implying that past performance will recur, a fair standard for reading any backtest.

Backtested investment returns and the SEC’s rules

Outside retirement planning, backtesting means applying a trading or allocation rule to past prices to see what it would have earned. It is easy to abuse: try enough rules and one will look brilliant on past data by luck alone.

Under the SEC’s marketing rule for investment advisers, performance backtested by applying a strategy to periods when it was not actually used counts as hypothetical performance. An adviser may show it only with policies designed to keep it relevant to the audience’s likely financial situation and objectives, and with enough information about its criteria, assumptions, risks and limitations. When a financial advisor shows you a backtested chart, ask when the strategy was designed, whether fees such as the expense ratio and taxes are included, and how it has done since it went live.

Illustrative numbers

Counting a Trinity-style test of a 4% inflation-adjusted withdrawal

Formula
Rolling periods = years of data − period length + 1; success rate = periods the money lasted ÷ rolling periods
Years of data
Calendar years of return history, such as 70 for 1926–1995
Period length
Length of the retirement being tested, such as 30 years
Periods the money lasted
Rolling periods that ended with money still in the portfolio

Rolling periods overlap, so they are not independent tests.

DataUS stocks, bonds and inflation, 1926–1995 (70 years)

Retirement length tested30 years

Rolling periods70 − 30 + 1 = 41, starting in 1926 through 1966

All stocks39 of 41 lasted (reported as 95%)

75% stocks, 25% bonds40 of 41 lasted (reported as 98%)

50% stocks, 50% bonds39 of 41 lasted (reported as 95%)

30-year periods with no overlap2

The rates look precise, but neighboring periods share most of their years, so one or two hard stretches decide the gap between 95% and 100%, and 70 years of data hold only two fully separate 30-year histories.

At a glance

Common backtesting pitfalls and how careful tests handle them

PitfallWhat goes wrongSafeguard
Overlapping periodsRolling windows share most of their years, so 41 results are not 41 independent testsCount separate periods and study the worst cases, not just the rate
Short, single-country recordA century of one market holds only a handful of deep crashes and inflation spellsAlso test lower assumed returns or reshuffled sequences
OverfittingTrying many rules and keeping the best one finds patterns that were luckFix the rule before testing, then check it on data it was not built from
Look-ahead biasThe test uses information no one had at the time, such as data revised after the factUse only data available at each decision date
Survivorship biasTesting only funds or stocks that still exist leaves out the failuresUse data that includes closed funds and dropped companies
Missing costs and taxesFees, trading costs and taxes lower real-world resultsSubtract fund fees and model taxes by account type

Put it in your plan

Backtesting in MoneyWhatIf

MoneyWhatIf’s Market Simulator is a single-path backtest of your whole plan. Choose an index, a historical start year and the plan year where that history lands, and the selected accounts grow by its actual calendar-year returns in their original order while cash keeps its own rate. The dialog names the worst year and the deepest drawdown, and the entire plan reruns, taxes and withdrawals included. Years before the landing and after the data ends use each account’s own return; history is never looped. Plan Resilience then reshuffles historical years into 100, 300 or 500 new orders.

Open your forecast

Common questions

Backtesting FAQs

Is historical backtesting reliable?

It is reliable as a record of what would have happened and weak as a forecast. A strategy that failed in history is a clear warning, because similar conditions could return. A strategy that survived only tells you it got through the specific sequences on record, and that record holds few truly independent long periods. Treat a backtest as a minimum stress test, not a guarantee.

Is historical backtesting better than Monte Carlo simulation?

Neither wins outright, because they fail in opposite ways. A backtest uses only sequences that really happened, so its crashes and inflation spells are real, but it cannot show anything worse than the record. A Monte Carlo simulation generates thousands of new sequences, including harsher ones, but they are only as realistic as its assumptions. A plan that holds up under both deserves more confidence than one tested either way alone.

Why do backtests of the 4% rule start in 1926?

Both Bengen’s 1994 paper and the Trinity study took their stock, bond and inflation figures from Ibbotson Associates’ Stocks, Bonds, Bills and Inflation yearbooks, whose annual series begin in 1926. Starting there gave them about 70 years of returns and inflation measured the same way, covering the Depression, World War II and the 1970s.

What is the difference between backtesting and forward testing?

A backtest applies a rule to data that existed when the rule was designed, so it may have been tuned, knowingly or not, to fit that history. A forward test fixes the rule first and tracks it on data that arrives afterward, often by paper trading. Holding back part of the history as an untouched out-of-sample test is a partial substitute. A strategy that passes both is far more convincing.