● 2014-15 NBA · SportVU tracking · 128,069 shots
How good
was that shot?
A model that puts a make probability on 128,069 tracked shots from the 2014-15 NBA season, using what the cameras saw: how far out the shooter was, how close the nearest defender got, the shot clock, and how long the ball was held.
It works, within limits. The best model reaches an AUC of 0.64: clearly better than guessing, far from certain. A lot of what decides a shot, like the shooter's touch, the quality of the contest, and plain luck, isn't in this data.
FG% by distance
all 128,069 shots- Restricted area63.6%+18.40–4 ft · 26,705
- Paint48.4%+3.24–8 ft · 23,311
- Short mid-range40.4%−4.88–16 ft · 19,812
- Long mid-range40.2%−5.016–22 ft · 23,251
- Three36.1%−9.122–26 ft · 31,060
- Deep three27.2%−18.026+ ft · 3,930
Rings are by distance only: the data has no x/y shot locations, so corners and wings are pooled. Colors and the small numbers compare each ring to the 45.2% league average.
- AUC, held-out shots
- 0.639
- 0.500 is a coin flip
- Log loss
- 0.648
- 5.9% better than always guessing 45%
- Brier score
- 0.229
- vs 0.248 for the same flat guess
- League FG%
- 45.2%
- 57,905 of 128,069 went in
Seven models, one ceiling
I compared a logistic regression, a random forest, a small neural net and two gradient-boosted tree libraries at near-default settings, then tuned the two boosters with a short random search. From best to worst, the whole field spans 0.0058 AUC (0.639 for LightGBM (tuned), 0.633 for XGBoost).
When models this different land this close together, the limit is what the features can tell you, not the algorithm. Tuning bought a little. Better information about each shot would buy more.
Test set · 25,614 held-out shots
Each line starts at the naive baseline, which predicts 45.2% for every shot.
ROC curves
How many makes the model catches as it calls more shots makes. The six other models sit almost exactly under the orange line.
Calibration
Test shots split into ten equal groups by predicted probability. On the diagonal, a 60% prediction goes in 60% of the time.
Why not accuracy?
With a 50% cutoff the model calls 62.0% of test shots correctly, against 54.8% for calling every shot a miss. But a yes/no call throws away the useful part. A 48% look and a 22% look are both “miss”, and the gap between them is what shot quality means. So the model is judged on its probabilities (log loss, Brier, calibration), and the probability is what the rest of this page uses.
What the model pays attention to
SHAP splits every prediction into pushes from each feature, toward a make or toward a miss. Averaged over 2,000 test shots, spacing and distance do most of the work. Who took the shot ranks eighth: once the model knows the situation, the shooter's own FG% adds little.
Units are log-odds. Near a 45% shot, +0.1 is worth about 2.5 percentage points.
Feature importance
Mean absolute SHAP value, tuned LightGBM
- Defender gap / shot distance0.257
- Shot distance0.205
- Touch time0.111
- Closest defender distance0.086
- Shot clock0.066
- Two or three0.037
- Dribbles0.021
- Shooter (encoded FG%)0.016
- Shooter's shot # in game0.015
- Game clock0.013
Top four, plotted in detail
How each top feature pushes
Each dot is one test shot. The orange line is the average push at that value.
Below about 0.4 (defender closer than 40% of the shot's length) this pushes toward a miss. Above it the push climbs fast: a defender as far away as the shot is long is worth roughly +0.8.
The defender paradox
Open shots aren't easy shots
Across all shots, defender distance barely correlates with a make (r = -0.001). Guarded within 2 ft, shots went in 45.3% of the time. With nobody within 6 ft, 44.7%.
Split by distance and the effect is obvious. At the rim, FG% climbs from 53% to 93% as the defender backs off. It hides in the total because open looks are mostly long ones: 58% of wide-open shots were threes, against 3% of tightly guarded ones.
That is why the engineered ratio, defender distance over shot distance, ranks first. It measures openness relative to how hard the shot already is.
Bars share one 0–100% scale. First and last numbers: tightest vs most open.
Who made more than they should have
A second LightGBM model is trained without knowing who took the shot. Its prediction is what an average player would make from the same distance, spacing and clock. Actual FG% minus that expected FG% is shot-making, the same idea as expected goals in soccer.
Every expected value is out-of-sample. The shots are split into five groups, and each group is scored by a model trained on the other four. Without knowing who shot, that model still reaches an AUC of 0.642, about the same as the main model's 0.639.
Kyle Korver leads: 49.2% on shots worth 38.2% for an average player, +11.0 points over 478 attempts.
Actual vs expected FG%
248 players with 200+ shots. Dot size is shot volume.
| Player | Over expectedOver expected, ±95% noise | |
|---|---|---|
| 1 | Kyle Korver49.2% vs 38.2% · 478 shots | +11.0 |
| 2 | James Johnson61.4% vs 51.6% · 311 shots | +9.8 |
| 3 | Alexis Ajinca59.7% vs 51.9% · 211 shots | +7.8 |
| 4 | DeAndre Jordan71.3% vs 63.7% · 393 shots | +7.5 |
| 5 | Chris Paul48.0% vs 41.0% · 885 shots | +7.0 |
| 6 | Stephen Curry48.5% vs 42.2% · 968 shots | +6.4 |
| 7 | Shaun Livingston52.5% vs 46.5% · 263 shots | +6.0 |
| 8 | Dirk Nowitzki46.2% vs 40.8% · 816 shots | +5.4 |
| 9 | Amar'e Stoudemire55.1% vs 49.7% · 336 shots | +5.4 |
| 10 | Beno Udrih48.9% vs 43.6% · 356 shots | +5.3 |
| 11 | Kevin Seraphin51.5% vs 46.5% · 336 shots | +5.0 |
| 12 | LeBron James48.9% vs 44.2% · 978 shots | +4.7 |
| 13 | Ed Davis60.3% vs 55.6% · 350 shots | +4.7 |
| 14 | Ben Gordon43.5% vs 38.9% · 255 shots | +4.6 |
| 15 | Cory Joseph50.6% vs 46.0% · 346 shots | +4.6 |
Dot: actual minus expected FG%, in percentage points. Whisker: the range luck alone would produce 95% of the time for that player's shot count. Scale runs from −16 to +16.
Most of the middle is noise
The whiskers show what luck alone does over a player's shot count. 46 of 248 players land outside them, against about 12 you'd expect by chance. So the ends of the list carry real signal, but most of the middle sits inside the noise band.
It measures making shots, not getting them
Expected FG% is built from each player's own attempts, so someone who creates easy looks is judged against an easy baseline. Shot creation is a separate skill this list doesn't credit.
The model can't see shot type
A dunk and a contested hook from 3 ft look alike to it. That is part of why centers show up at both ends of the list.
Price a shot
Set up a shot and read off the model's make probability. There's no server here: every number comes from predictions made ahead of time across the five main inputs and shipped with the page.
Make probability
–
Start from a common look
Threes start at 22 ft in the corners
6+ ft counts as open
4 s or less is late clock
How long the shooter held the ball
0 = catch and shoot
Distance × defender map
Every distance and defender gap at the current clock, touch and dribbles. Click the map to move the shot.
Predictions come from the tuned LightGBM model on a grid of about 207k points and are blended linearly in between. Tree models move in steps, so small slider moves can jump. Hatched cells had fewer than 25 real shots, so treat them as extrapolation.
How it was built
The notebook is the full record: data cleaning, EDA, PCA and t-SNE, five baselines, tuning, SHAP, and the player analysis. A Python package reproduces it, and one export script writes every number on this page. One deliberate change: the notebook scored most leaderboard shots in-sample, while this page uses out-of-fold predictions. Player numbers moved by about a tenth of a point on average.
01 Data
128,069 shots from 904 games, October 2014 to early March 2015, from the Kaggle NBA shot logs (SportVU tracking via NBA.com). Each row has shot distance, closest defender and their distance, shot clock, dribbles, touch time, and the result.
02 Cleaning
Game clock parsed to seconds. 5,567 missing shot clocks are flagged, then filled with the game clock (capped at 24), since the shot clock switches off late in periods. 312 negative touch times, a sensor glitch, are clipped to zero.
03 Features
24 inputs: the raw tracking values, log dribbles and touch time, six distance zones, flags for catch-and-shoot, late clock (4 s or less), open (6+ ft), shot clock off, home and buzzer beaters, and defender distance divided by shot distance plus one.
04 No leakage
Shooter and defender IDs become smoothed FG% values, (n·mean + 50·league) / (n + 50), computed out-of-fold on the training set so no shot ever sees its own result.
05 Training
Stratified 80/20 split (102,455 / 25,614). Five model families at near-default settings, then a six-config random search for XGBoost and LightGBM, scored on 3-fold cross-validated log loss.
06 Scoring
Log loss and Brier score first, since the output is a probability. AUC for ranking, calibration curves to check the probabilities mean what they say. Accuracy is reported but isn't the target.
Limits
Shot quality is noisy
Project the shots onto their first two principal components (together 48% of the variance) and made and missed shots overlap almost everywhere. Some regions lean one way, since close shots go in more often, but there is no boundary to find. The data groups shots by type and situation, not by outcome. That is the ceiling every model on this page runs into.
- AUC 0.64 is modest. The model sorts good looks from bad ones, but any single shot is still close to a coin flip.
- No shot location beyond distance, no shot type, no play type, no help defense, no fatigue beyond the period and game clock.
- One partial season, 2014-15. The league has moved further toward threes and rim attempts since.
- Closest defender is measured at the moment of the shot. A late closeout and a defender who never left look the same.
2,000 random shots, PC1 vs PC2, standardized tracking features.
Reproduce it
The export trains everything once in about 30 seconds and checks each metric against the notebook. Library versions are pinned. See the notebook and the export script.
$ pip install -r requirements.txt
$ python scripts/export_site_data.py
$ cd web && npm install && npm run build