Wrangling NBA Play-by-Play with hoopR
My first shot chart came out shaped like a butterfly. Two mirrored blobs with a hole in the middle. The bug turned out to be a fact about basketball courts I had never thought about.
The whole data layer of Hoop Vision is one R
script. It pulls a season from hoopR, cleans it, does every slow
calculation once, and writes small JSON files the app reads. Nothing in the
app itself ever touches the network.
Getting a season is three lines:
library(hoopR)
pbp <- load_nba_pbp(seasons = 2026) # hoopR labels a season by the year it ENDS
nrow(pbp)
#> [1] 642472
642,472 rows. Every substitution, rebound, timeout and shot from all 1,326 games of 2025–26. It downloads in about four seconds because hoopR reads a prebuilt file off a GitHub release rather than scraping.
That was the last easy thing that happened.
The butterfly
I filtered to shooting plays, grabbed coordinate_x and
coordinate_y, and plotted them. I got two dense clouds at
opposite ends of a long rectangle with empty space between.
I stared at it for a while, and when it clicked I felt genuinely stupid: teams shoot at both baskets. They switch at halftime. ESPN records where the shot happened on the actual floor, so a team's attempts are scattered across both ends depending on the half. Of course they are. That is how basketball works.
range(shots$coordinate_x) #> -46.75 46.75 <- length of the floor
range(shots$coordinate_y) #> -25.00 25.00 <- width of the floor
So coordinate_x is not left-to-right, it is
baseline-to-baseline. Both of my assumptions about the axes were wrong at
once.
The fix folds both halves onto a single basket. Distance from the nearest
baseline is 46.75 - abs(x), and the sideways position is
y, sign-flipped on one half so left and right stay consistent
after the rotation.
shots <- pbp |>
filter(shooting_play) |>
transmute(x = coordinate_x, y = coordinate_y)
What convinced me the fold was right was not the picture. It was that the rim landed at 5.25 feet from the baseline — exactly where a real NBA hoop is. I did not put that number in. It fell out of the arithmetic. That is the best feeling in this whole project: when the data agrees with the physical world about something you never told it.
The bug that assertion didn't catch
The geometry check passed and I felt good about myself for about a week. Then I looked at the finished chart and the three-point line was cold. Not slightly — the entire arc was returning about 0.88 points per shot, which would mean the league shot 29% from three. That is not a number that happens.
The problem was how I decided whether a shot was worth two or three:
value = ifelse(score_value >= 3, 3L, 2L) # wrong
score_value is the points a play scored. On a miss it
is zero. So every missed three fell through to the else and got
labelled a two. Made threes counted as threes, missed threes counted as
missed twos, and the three-point zones got dragged down by tens of thousands
of misattributed misses.
It is a nasty bug because the code does not look wrong, the script does not error, and the chart still renders a perfectly plausible court. It is only wrong if you know what the number should be.
The column I wanted was sitting right there:
pbp |> count(points_attempted)
#> points_attempted n
#> 1 2 137557
#> 2 3 98557
Populated on every row, made or missed. So now it is just
value = as.integer(points_attempted), with a second assertion
guarding it — and this one tests against reality rather than against my own
reading of the code:
# League-wide this is ~55% on twos and ~36% on threes every single year.
stopifnot(
between(fg_pct_2, 0.48, 0.62),
between(fg_pct_3, 0.31, 0.40)
)
The build now prints 2PT 55.0% (1.10 pps) | 3PT 35.7% (1.07 pps)
on every run, and refuses to write the files if it ever doesn't.
My geometry assertion tested the thing I had just fixed. It did not test the thing I hadn't thought of yet. Now when I find a bug I try to write the check for the category of mistake, not the specific one.
Two more things that were hiding
My standings table came out with 33 teams. Three of them were STARS, STRIPES and WORLD — the All-Star weekend rosters, filed by ESPN as if they were franchises. I filter them out by games played rather than by name, since real teams play 82 and exhibition rosters play two or three.
Then San Antonio and New York had 83 games. That one is the NBA Cup final: ESPN tags it as regular season, but it does not count in the standings. I could have hardcoded the game id, and nearly did. Instead I pull the schedule, which labels it directly, so the fix still works next season when it is two different teams:
Hexagons, and why not squares
Once the coordinates were right I binned them. Square bins make a shot chart look like a spreadsheet and produce grid artifacts that are not in the data. Hexagons tile without those seams, and every neighbour sits the same distance away.
Converting a point to a hex is a coordinate transform plus rounding, and the rounding is the part that is easy to get wrong. You cannot round the two axial coordinates independently, because the axes are not perpendicular and that misassigns points near the hex edges. You convert to cube coordinates, round all three, then repair whichever moved furthest.
219,607 shots go in. 243 hexagons come out, each carrying its attempt count, field goal percentage and points per shot:
Why it all gets written to disk
Halfway through building this, a hoopR call failed on me with a
503 from GitHub — mid-pull, no warning. If that had happened while I was
showing the app to someone, the demo would simply have been over.
So the script caches the raw download locally, every run after the first is completely offline, and the app never calls an API at all. It reads five JSON files totalling under 400 KB. The slow, fragile part happens once, on my machine, where a failure costs me thirty seconds instead of a meeting.
Every field is documented in DATA.md, and the full script is in the repo.