How We Built a Model That Hits 75.4% on Its Best Picks
Behind the scenes of the PuckCast NHL prediction model. How we built a system that hits 75.4% on A-grade picks across 6,560 tested games.
Everyone has a hockey opinion. Not everyone has a model.
When we started building PuckCast, we weren't trying to be a prediction site. We were trying to answer a question: is it actually possible to predict NHL games with any real consistency, or is it mostly noise?
After a full season of testing, here's what we found out — and why the answer is more nuanced than most people expect.
The Problem: Why NHL Is Hard to Predict
Hockey is the hardest major North American sport to predict. That's not a complaint — it's a design feature. The NHL is built around parity. The salary cap compresses talent distribution. The playoff format is a best-of-seven coin flip factory. And more than any other sport, single games are subject to enormous variance.
Three specific factors make it brutal:
Goaltending variance. A goalie can steal a game. Any goalie, on any night, against any team. A .920 save percentage is excellent; a .940 is miraculous; and both can happen in the same week for the same goalie. This is the biggest single source of prediction noise in the sport.
Shot volume ≠ shot quality. A team can generate 40 shots and still lose because those shots came from bad angles on a locked-in goalie. Another team can win with 22 shots because all 22 were from the slot. Traditional shot-based statistics miss this almost entirely.
Schedule and fatigue. Teams play 82 games in roughly 180 days. Rest, travel, back-to-backs, and altitude all matter. A team playing their third game in four nights on the road against a rested opponent is a very different bet than their raw record suggests.
We built PuckCast to account for all three of these factors and more.
The 8 Factors We Use
Our model pulls data on eight core input categories before generating any prediction:
1. Recent Form — Rolling 10-game performance, weighted so the most recent games count more. A team's last three games matter more than their record six weeks ago.
2. Home/Away Splits — Home ice advantage is real in hockey, but it varies by team. Some franchises are dramatically better at home; others are nearly indifferent. We model each team's actual home/away differential rather than applying a league-wide constant.
3. Goaltending Performance — Quality starts, goals saved above expected (GSAx), high-danger save percentage, and starter vs. backup projections. We also track goalie fatigue and days of rest.
4. Rest and Schedule Density — Back-to-back games, travel distance, days since last game for both teams. Fatigue effects are real and measurable.
5. Head-to-Head History — Recent H2H results and underlying metrics, weighted by recency. Some matchups have genuine systematic advantages that persist over time.
6. Shot Quality and Expected Goals — Corsi and Fenwick as baselines, but we go deeper into shot location, traffic in front of the net, and expected goals (xG) per shot attempt. This is where we diverge most significantly from simple shot-based models.
7. Power Play and Penalty Kill Efficiency — Special teams account for a real percentage of goals. Teams with elite PPs and PKs are systematically undervalued in naive models.
8. Opponent Strength — Not just record, but quality-adjusted performance. Beating bad teams doesn't mean what it sounds like; losing to elite teams in close games isn't as damning as the box score suggests.
How Confidence Grades Work
Every game prediction comes with a confidence grade: A, B, or C.
Grade A (>20 pts edge) — Our model has strong conviction. High-confidence signals align across multiple factors. These are the picks where the model isn't just leaning one direction, it's leaning hard. A-grade picks hit 75.4% across 6,560 walk-forward tested games.
Grade B (5-20 pts edge) — Solid lean, but one or more factors are uncertain or conflicting. Still valuable at 61.0% historically, but more variance in outcomes.
Grade C (0-5 pts edge) — Lean exists, but the model is flagging elevated uncertainty. Goalie uncertainty, schedule conflicts, or contradictory recent form. 53.5% historically. Informational rather than high-confidence.
We publish the full record on our track record page so you can verify it yourself.
What We're Still Improving
We're honest about where the model falls short, because that's how you make it better.
Goalie projections on non-starter situations are still our weakest area. When a backup comes in mid-game or a starter is a game-time decision, the model's confidence drops significantly. We're building better starter probability estimates, but this is genuinely hard.
In-season adjustments to team identity. A coaching change, a major trade, a system shift — these can fundamentally change how a team plays, and retrospective data doesn't capture it fast enough. We're working on a weighting system that discounts pre-change data more aggressively.
Momentum effects beyond rolling windows. There's evidence that teams on 5+ game winning streaks perform above their metrics for non-random reasons. We're still evaluating whether this is real signal or small-sample noise.
The goal isn't a perfect model. Perfect doesn't exist in a sport this variable. The goal is a model that's right more often than it's wrong on the bets it's most confident about — and we're achieving that.
Follow @PuckcastAI for daily picks, model updates, and results tracking.
Want to dig deeper? Check out the full model accuracy breakdown and our methodology page for the technical details.
All accuracy figures are tracked live and reflect real game results against model predictions. Nothing is retroactively adjusted. The record is the record.