PlayDecoded
Menu

Our Methodology

How we calculate win probabilities—and why you should trust the numbers

By PlayDecoded Analytics Team·Updated 2026-02-27

Our Approach

Sports prediction is fundamentally a problem of incomplete information. You can't know everything that affects the outcome of a game—locker room dynamics, a player's sleep quality, whether the ref had a bad morning. What you can do is measure what's measurable and weight it based on how predictive it's actually been.

That's what we do. For every game we score a set of measurable factors (team strength, home and road form, rest, weather, injuries and more) and combine them into a single win probability. The number isn't a prediction; it's our best estimate of each team's chances given what we know right now.

We show our work because you should be able to see what's driving the number. If you think we're underweighting a factor, that's useful information for you.

The Models

Elo Ratings

Every team starts with a baseline rating. Win and you gain points; lose and you drop. The amount transferred depends on the expected outcome—beating a strong team is worth more than beating a weak one. This system, originally designed for chess, has proven remarkably effective across sports because it captures relative team strength without overreacting to single results. Between seasons each rating keeps about two-thirds of its distance from average, so last year still counts but doesn't decide this year.

Based on: Elo (1978), Bradley & Terry (1952)

From Factors to a Probability

Each factor gives the home team a score: positive when it favors them, negative when it favors the visitors. We multiply each score by that factor's weight, add them up with a small built-in home edge, and pass the total through a logistic (S-shaped) curve that turns it into a win probability between 5% and 95%. The aim is calibration: when we say a team has a 70% chance, those teams should win about 70% of the time. We don't just assert that; we track it by confidence band on our accuracy page, counting only picks made before kickoff.

Factor Weighting

Not all factors matter equally, and what matters varies by sport. Each sport has its own set of weights: in MLB, runs scored and allowed carry the most weight; in the NHL, goals allowed does; weather only counts for open-air NFL and MLB venues, and back-to-backs only matter in the daily-schedule sports. The weights are set per sport by us, not learned automatically, and they change only when we release a new model version. Every game page shows what each factor contributed.

Based on: Schwartz & Barsky (1977) on home advantage, Massey (1997) on sports ratings

Academic Foundations

Our methodology draws on decades of research in sports analytics, statistics, and decision science. Here are the foundational works that inform our approach:

  • Elo, A. E. (1978). The Rating of Chess Players, Past and Present. Arco Publishing. Link
  • Bradley, R. A., & Terry, M. E. (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4), 324–345. doi:10.2307/2334029
  • Schwartz, B., & Barsky, S. F. (1977). The Home Advantage. Social Forces, 55(3), 641–661. doi:10.2307/2577461
  • Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3. doi:10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2
  • Silver, N. (2012). The Signal and the Noise: Why So Many Predictions Fail—but Some Don't. Penguin Press. Link
  • Massey, K. (1997). Statistical Models Applied to the Rating of Sports Teams. Bluefield College. Link

Data Sources

Good models need good data. Here is where ours comes from and how often it refreshes:

Game Data

Schedules, scores and team stats come from ESPN's public sports data feeds, synced every 30 minutes. A daily check settles any game still marked scheduled a day after its start from ESPN's final result.

Injury Reports

ESPN's league injury reports, synced every 3 hours and every 15 minutes for games starting within 3 hours. Each player's listed status (out, doubtful, questionable) counts, weighted by position.

Weather Data

OpenWeatherMap daily forecasts for open-air NFL and MLB venues, pulled once a day at 06:00 UTC. Temperature and wind count, mainly in cold games; domes and retractable roofs are treated as no weather effect.

Validation

A pick only counts toward accuracy if it was made before the game started, and each model's record is kept separate. Factors with missing data score zero and say so on the game page instead of guessing.

Model Accuracy

We hold ourselves accountable. Here's how we measure whether our probabilities are actually accurate:

~53%
Close-call outright accuracy
Near toss-ups (favorite under 60%)
~84%
High-confidence outright accuracy
75%+ probability calls
~56%
Overall outright accuracy
Across all games and sports

Figures cover the live factor model only; the retired XGBoost model's record is shown separately. See live accuracy data →

Accuracy glossary

Outright accuracy
The share of games where the team we gave over 50% actually won, ties excluded. This is the number on our accuracy page.
Against the spread (ATS)
Whether a pick beats the sportsbook point spread. Our model predicts winners, not margins, so we do not publish an ATS record.
Brier score
The average squared gap between our probability and the result, where 0 is perfect and 0.25 is a coin flip. We do not publish one yet.
Calibration
Whether teams we give 70% win about 70% of the time. The chart on the accuracy page checks this for every probability band.

Calibration

Calibration asks a simple question: when we say a team has a 70% chance, do those teams actually win about 70% of the time? We test it by bucketing every prediction and comparing predicted vs. actual win rates. A perfectly calibrated model traces the diagonal. Ours holds up well on near toss-ups and on our most confident calls, and runs a little overconfident in the middle bands, which is exactly what the accuracy page shows in full.

Brier Score

The Brier score measures the accuracy of probabilistic predictions—it's the mean squared error between predicted probabilities and actual outcomes (0 or 1). Lower is better. A score of 0 means perfect predictions; 0.25 is no better than flipping a coin. We do not publish our Brier score yet; the calibration chart is the probability check we show today.

Based on: Brier (1950)

Win Rate by Confidence

We bucket games by how confident we were. Near toss-ups land close to a coin flip, as they should. The more confident we are, the higher our hit rate climbs, and our 75%+ calls are where we're strongest. The live figures for each band are in the cards above and on our accuracy page.

What We Can't Measure

No model captures everything. Revenge games, contract years, coaching changes mid-season, locker room drama—these matter but resist quantification. We also can't predict surprise inactives until they're announced. A 35% underdog still wins 35% of the time; that's not an error, it's uncertainty. Use our numbers as one input among many.

See the methodology in action

Every game page shows the factor breakdown driving our probabilities.

Learn More

Learn about specific aspects of our analysis: