To backtest a betting strategy, you write the rule down exactly and run it over past races or games, using only the prices and information available when each bet would have gone on. Then you judge it on data it was never tuned on. Skip any of those steps and the test measures your hindsight as well as the rule.
With illustrative figures, a rule tuned on two seasons of racing showed a +6.5% yield over 2,400 bets, then -3.1% over the next 1,150 bets it had never seen. A sound betting strategy finds that out on past data, before a dollar is staked.
The data a backtest needs
A backtest is only as honest as its prices. For every bet the rule would have made, you need:
| Field | Why it matters |
|---|---|
| The price on offer when the rule fires, with its time stamp | The rule bets then, not at the open, the close or the best price of the day |
| Who offered it | A price from a bookmaker you could not use is not one you could have taken |
| The closing price, such as the Betfair Starting Price (BSP) in racing | It is the benchmark for closing line value |
| Results, scratchings, voids and dead heats | They decide what each bet returned |
| Stewards' deductions on fixed-odds racing bets | A late scratching cuts what winning bets placed before it are paid |
| Commission and other costs | They come out of exchange winnings |
When a runner is scratched late, the race-day Stewards declare deductions, in cents in the dollar, on winning fixed-odds bets placed before the scratching. Racing NSW's Rules of Betting (BR 14) set this out for that state's on-course bookmakers, and your bookmaker's terms say how it applies them. A backtest that pays every winner in full overstates those races.
Historical odds data covers where past prices come from and each source's gaps. If none records prices when your rule fires, the cleanest data is your own snapshots, saved from now on at the time you would bet.
Write the rule down before you run it
A rule a computer could follow leaves no room to bend it after seeing the results. "Back runners that look well placed" is not a rule. "Back any runner whose best bookmaker price 10 minutes before the scheduled jump is at least 5% above the exchange back price at that moment, between $2.50 and $10.00, for $20" is one. The timing, the price limits and the stake each change the result, so fix them all.
Date the rule before the first run. Every change after that is a new version, and the number of versions you try decides how much a good result is worth. The same steps test a betting system someone else sells: get its rule in writing and run it yourself on prices it was not built on.
In-sample and out-of-sample testing
Split the history by time, never at random. Tune the rule on the earlier part (in-sample), then run the finished rule once on the later part (out-of-sample). A random split mixes seasons, so lessons from later races leak into the tuning. Two illustrative racing rules, tuned on seasons 1 and 2 and tested on season 3:
| Rule | Seasons 1 and 2: bets | Yield | Average CLV | Season 3: bets | Yield | Average CLV |
|---|---|---|---|---|---|---|
| A: a form filter, the best of 12 versions tried | 2,400 | +6.5% | +0.4% | 1,150 | -3.1% | -0.6% |
| B: a price rule against the exchange, one version | 2,900 | +2.9% | +3.6% | 1,300 | +1.8% | +3.2% |
Rule A looks better in-sample, and luck explains why that means little. At its average price of $5.00, one standard error of the yield over 2,400 bets is the square root of (5.00 - 1) / the square root of 2,400 = 2 / 49.0 = 4.1 points. Twelve versions with no edge would, if independent, give a best result about 1.63 standard errors above zero on average: 1.63 x 4.1 = 6.7 points. Rule A's +6.5% is about what keeping the luckiest of 12 produces, and its average CLV, near zero in both periods, says the same.
Rule B's yield proves nothing either: at an average of $5.40, +1.8% over 1,300 bets sits well inside one standard error of 5.8 points. What held was its CLV, above +3% in both periods, and sample size in betting shows why closing line value settles so much sooner than profit.
Overfitting, the same trap inside a model, is covered in how to build a betting model.
Risk: Betting involves risk. Past results are no guarantee of future results, and an edge that held in a test can vanish live. See responsible gambling for limits and support.
Look-ahead bias: using what you could not have known
Look-ahead bias is any use of information that did not exist when the bet would have gone on. In Australian racing data it hides in four places:
- The bet price. Test at the top fluc, the highest price during betting, and you assume you always caught the best price of the day. With illustrative prices, a $10 winner at a $6.50 top fluc returns $65, while the $5.60 on offer when your rule fired returns $56: $9 that never existed, on one bet. Against a $5.20 BSP, the same bet's CLV jumps from 5.60 / 5.20 - 1 = +7.7% to 6.50 / 5.20 - 1 = +25.0%.
- The close. Testing at the BSP or the final fixed price, when the rule fires 10 minutes earlier, borrows the market's last word.
- The field. Filters such as "fields of 10 or fewer" or "the favourite" read from the final field or market, after late scratchings and moves.
- The ratings. A track rating upgraded on race day, or ratings averaged over runs that came later, feeds the future into the past.
In sport, confirmed line-ups and late team news belong only in tests whose bets go on after the news broke.
Survivorship bias: testing only what survived
Survivorship bias is judging a method by the cases still standing. In a backtest it creeps in when the data, or the rule, is chosen with hindsight:
- Bookmakers that have since closed or rebranded drop out of price histories, so today's list of bookmakers misstates the best price on offer back then.
- Races and games with incomplete data get deleted, and abandoned meetings, voids and refunds go with them.
- Runners scratched after your rule would have fired vanish from the data, along with the refund or deduction your bet would have met.
- In soccer, a history built from this season's clubs drops the relegated sides your rule would have bet on.
- A rule chosen because it is popular now is itself a survivor: the systems that failed stopped being sold.
Measuring a backtest with yield and closing line value
Report three numbers for every test, each beside the number of bets:
- yield = profit after commission and deductions / total staked, which records often label ROI (ROI vs yield)
- CLV per bet = the price the rule took / the close - 1, using the BSP in racing and a margin-free close in sport
- standard error of the yield = about the square root of (average price - 1) / the square root of the bet count
Then allow for what a backtest leaves out: every backtest bet is accepted in full at the recorded price. Live, some prices are gone before the bet lands and some bets are refused or cut. A backtest edge smaller than those costs is no edge at all. Paper betting tests the rule on live prices without money, but only real bets show refusals, cut stakes and the prices you get.
Risk: Betting involves risk. A backtest describes the past with every bet accepted, part of any backtest profit is luck and tuning, and there is no guarantee of profit. See responsible gambling for limits and support.
From backtest to small live stakes
A live trial needs its pass and stop rules written down before the first bet, because yield over a few hundred live bets is mostly noise. If Rule B's true yield were about 2%, 300 live bets at an average of $5.40 would carry a standard error of the square root of 4.4 / the square root of 300 = 12.1 points. Anything from about -22% to +26% would fit the backtest, so yield cannot pass or fail the trial. Judge it on what settles sooner, against thresholds set in advance (illustrative here):
| Measure | In the backtest | Pass mark after 300 live bets |
|---|---|---|
| Average CLV | +3.2% | Above +1.5% |
| Fill rate: signals bet at or above the rule's price | 100%, by assumption | Above 75% |
| Average price gap: live price / signal price - 1 | 0% | Better than -2% |
Write the stop rules beside the pass marks. For example, stop after 100 bets if fewer than 60% of signals were filled, stop after 200 if average CLV is below zero, and stop at a loss limit sized from your bankroll and the backtest's largest drawdown. Keep the stake fixed for the whole trial. A losing run inside it is a reason to check the rule, never to raise the stake and win it back.
Put a deposit limit on each bookmaker account before the first bet. Under the National Consumer Protection Framework, online bookmakers must offer one, a decrease applies at once and an increase only after 7 days, so the cap holds on a bad day.
Where B337 fits once a rule becomes code
API Odds delivers live racing and sports odds from Australian bookmakers and Betfair over the B337 API and places no bets. The Execution API places racing and sports bets from your own code into sessions running on your computer, in bookmaker accounts you already hold, and that computer has to be on, awake, online and running B337. The target price your code sends with a bet is a floor: below it the bet is refused, and at or above it the bet goes on at the bookmaker's current price, so send the lowest price your rule accepts.
The API places what your code sends and picks no selections. Both products are set up with the team, with the price confirmed before you pay, coding is required, and bets placed through B337 use credits. The betting API shows how a bet travels from your code to a session. Many bookmakers restrict automated betting in their terms, and automation can bet far more, far faster, than by hand, so keep those deposit limits in place. Bookmaker and exchange names are trade marks of their owners. B337 is not affiliated with them.
For free and confidential support call 1800 858 858 or visit gamblinghelponline.org.au.
Risk: Betting involves risk. Code places exactly what it is told, mistakes in the rule included, and a bookmaker can restrict an account or void bets. See responsible gambling for limits and support.