To build a betting model, you pick one thing to predict, collect the data known before each game or race, fit a simple model and turn its outputs into fair prices. Then you test them on games the model never saw and against the closing line. A model earns its keep only when its prices beat the market's after the margin, so most of the work is testing, not building.
With illustrative numbers, a ratings model makes the home side a 14-point favourite in an AFL game, with a 37-point standard deviation in its past errors. That gives the home side a 64.7% chance: a fair head to head price of $1.54, and a fair $1.82 at -9.5 on the line.
Step 1: choose what the model predicts
Pick one target, in a market you will actually bet:
| Target | What it prices | Australian example | What the model outputs |
|---|---|---|---|
| Win probability | Head to head and racing win markets | NRL head to head; a race's win market | A chance for each side or runner, adding to 100% |
| Margin | Line betting and margin bands | AFL line | An expected margin and how widely results spread around it |
| Total | Overs and unders | AFL total points; soccer goals | An expected total and its spread, or a scoring rate for each side |
A margin model is a good first build for AFL and NRL, because one number prices head to head, line and margin bands together. Racing is a win-probability problem across a whole field, and framing a betting market turns ratings into chances that add to 100%. For low-scoring sports, a Poisson model for goals turns each side's expected goals into a chance for every score.
Step 2: data and features known before the bet
The data is several seasons of results with scores, plus what was known before each game: venue, travel, days since the last game and team changes. For racing it is barrier, weight, class, distance, track condition and recent runs. Keep the prices on offer at the time as well, each with a time stamp, because testing needs them; historical odds data covers where to get them.
Features are the inputs built from that data, such as a rating difference, home ground and rest days. Start with a handful: each extra one is another chance to fit noise.
An online sports bet has to be placed and accepted before the event begins, with a provider licensed in a state or territory (Interactive Gambling Act, s 10B, s 8A(3) and s 15AA). So a sports model works from what is known before the start. A feature built from later information, such as a team change after your bet would have gone on, a late scratching, or the closing price, leaks the answer into the test.
Step 3: a simple first model you can check by hand
A ratings model is enough to start. Give each team a rating in points, and:
- predicted margin = home rating - away rating + home ground advantage
- after each game, new rating = old rating + k x (actual margin - predicted margin) for the home side, with the away side moved the same amount the other way
Example: Illustrative numbers, not real teams. The home team is rated +6, the away team -4, and home ground is worth 4 points, so the predicted margin = 6 - (-4) + 4 = +14. With k = 0.1, a 30-point home win lifts the home rating by 0.1 x (30 - 14) = 1.6 points and drops the away rating by 1.6.
To turn a margin into chances you need the spread of the model's errors: the standard deviation of (actual margin - predicted margin) over past games. Say it is 37 points (illustrative: measure your own). Treating results as normally spread around the prediction, and ignoring the small chance of a draw, a spreadsheet's NORM.DIST does the rest:
| Market | Calculation | Chance | Fair price: 1 / chance |
|---|---|---|---|
| Home head to head | NORM.DIST(14, 0, 37, TRUE) = 0.6474 | 64.7% | $1.54 |
| Away head to head | 1 - 0.6474 | 35.3% | $2.84 |
| Home -9.5 on the line | 1 - NORM.DIST(9.5, 14, 37, TRUE) | 54.8% | $1.82 |
A machine learning model fits its own formula to many features, at the cost of more data and more ways to overfit, as machine learning in betting explains.
Step 4: turn outputs into prices and compare with the market
Fair price = 1 / chance, and the comparison that matters is with the market after its margin is taken out, not with one bookmaker's raw price. With illustrative prices of $1.61 for the home side and $2.32 for the away side:
- implied chances: 1 / 1.61 = 62.1% and 1 / 2.32 = 43.1%, a book of 105.2%
- margin-free home chance: 62.1 / 105.2 = 59.0%
- your model's 64.7% at $1.61: 0.6474 x 1.61 - 1 = +4.2%
A gap of 5.7 points against the whole market usually means the model is missing something the market knows, such as team news. A common answer is to blend your chance with the margin-free market chance and test the weight. At 40% model and 60% market (illustrative weights), the chance is 0.4 x 0.6474 + 0.6 x 0.590 = 61.3%, worth 0.613 x 1.61 - 1 = -1.3%: no bet. Removing the bookmaker margin compares the ways to get the market's chance, and the fair odds calculator applies them.
Risk: Betting involves risk. A model that disagrees sharply with the market is often missing information, and a bet with a real edge can still lose, often. See responsible gambling for limits and support.
Step 5: test the model on games it never saw
Fit on earlier seasons and test on the latest, run as if live: move forward a round at a time, refit only on games already played, and use only prices on offer when each bet would have gone on. Score the chances, not the winners picked. The Brier score is the average of (forecast - result) squared, with 1 for a win and 0 for a loss, and lower is better. Score the market's margin-free chances on the same games, because a useful model has to beat the market, not a coin toss.
With five illustrative test games:
| Game | Model's home chance | Market's home chance | Home won | Model's score | Market's score |
|---|---|---|---|---|---|
| A | 72% | 64% | Yes | 0.0784 | 0.1296 |
| B | 68% | 58% | No | 0.4624 | 0.3364 |
| C | 55% | 52% | Yes | 0.2025 | 0.2304 |
| D | 30% | 38% | No | 0.0900 | 0.1444 |
| E | 60% | 50% | No | 0.3600 | 0.2500 |
| Average | 0.239 | 0.218 |
Game B's model score is (0.68 - 0) squared = 0.4624. The model called three of five games correctly, but it was more confident on the two it got wrong, so it scored worse than the market. Five games prove nothing: run the same test over a whole season. Backtesting betting strategies sets out the full method, including the traps in historical prices.
Step 6: test against the closing line
Profit is slow to judge, as sample size in betting explains. At two standard errors, telling a 4% edge at $1.90 apart from luck takes about 4 x (odds - 1) / (edge x edge) = 4 x 0.90 / (0.04 x 0.04) = 2,250 bets. Closing line value answers sooner: for every bet the model would have made, record the price taken and the margin-free closing price, then CLV = price taken / margin-free close - 1.
Example: Illustrative prices. The model backs the home side at -9.5 on the line at $1.90. The line closes at $1.75 for the home side and $2.08 for the away side: closing book = 1 / 1.75 + 1 / 2.08 = 1.0522, margin-free close = 1.75 x 1.0522 = $1.841, and CLV = 1.90 / 1.841 - 1 = +3.2%.
An average CLV that stays above zero over hundreds of bets is the earliest sign a model knows something the market learns later. Good backtest profit with negative CLV usually means the backtest was lucky or leaky. Closing line value covers the method.
Overfitting and other traps
| Trap | What it looks like | The check |
|---|---|---|
| Overfitting | Strong on the seasons it was fitted to, ordinary on the next | Keep a test season untouched, and prefer fewer features |
| Testing many ideas on one data set | Try 40 features or filters and keep those passing a 5% significance test: about 2 pass by luck alone (40 x 0.05), and at least one passes 87% of the time (1 - 0.95 to the power 40) | Write the idea down before testing, and confirm it on new data |
| Leakage | Inputs from after the bet: late team changes, scratchings, the closing price | Time-stamp every input and ask whether you could have known it when betting |
| Ignoring the margin | Judging the model against raw implied chances, which add up to more than 100% | Compare with margin-free chances, as in step 4 |
| Small samples at long prices | A racing model betting $8.00 to $15.00 chances swings hard for hundreds of bets | Judge it on CLV before profit |
| Survivorship | Testing only on teams, tracks or rules still around today | Keep everything that existed at the time; see survivorship bias |
A model is one route to the fair price a betting strategy rests on, and every route needs the same testing.
Where B337 sits beside your own model
B337 has no model of its own and makes no selections. Two of its products sit at either end of yours:
- API Odds sends prices from Australian bookmakers and Betfair to your code over the B337 API: the prices your model has to beat. It places no bets, includes Terminal View and needs coding.
- The Execution API, set out on the betting API page, passes each bet your code sends to a session on your computer, logged in to a bookmaker account you hold. That computer has to be on, awake, online and running B337, or nothing is placed. Your model's lowest acceptable price goes in as the target, a floor below which the bet is refused.
Both are set up with the team, and the price, including any charges on top of the plan price, is confirmed before you pay. Bets placed through B337 use credits. Many bookmakers restrict or prohibit automated betting and third-party access in their terms, so using B337 may breach them, and a bookmaker can limit stakes, void bets or close an account; that risk is yours. Bookmaker and exchange names are trade marks of their owners. B337 is not affiliated with them.
For free and confidential support call 1800 858 858 or visit gamblinghelponline.org.au.
Risk: Betting involves risk. An edge that held up in testing can shrink once the market prices in what the model knows, and automation can bet far more, far faster, than by hand, so keep a deposit limit on every account. There is no guarantee of profit. See responsible gambling for limits and support.