← All writing

Wrangling NBA Play-by-Play with hoopR

Data Pipeline 9 min read• 642,472 rows• R / hoopR

My first shot chart came out shaped like a butterfly. Two mirrored blobs with a hole in the middle. The bug turned out to be a fact about basketball courts I had never thought about.

The whole data layer of Hoop Vision is one R script. It pulls a season from hoopR, cleans it, does every slow calculation once, and writes small JSON files the app reads. Nothing in the app itself ever touches the network.

Getting a season is three lines:

library(hoopR)

pbp <- load_nba_pbp(seasons = 2026)   # hoopR labels a season by the year it ENDS
nrow(pbp)
#> [1] 642472

642,472 rows. Every substitution, rebound, timeout and shot from all 1,326 games of 2025–26. It downloads in about four seconds because hoopR reads a prebuilt file off a GitHub release rather than scraping.

That was the last easy thing that happened.

The butterfly

I filtered to shooting plays, grabbed coordinate_x and coordinate_y, and plotted them. I got two dense clouds at opposite ends of a long rectangle with empty space between.

I stared at it for a while, and when it clicked I felt genuinely stupid: teams shoot at both baskets. They switch at halftime. ESPN records where the shot happened on the actual floor, so a team's attempts are scattered across both ends depending on the half. Of course they are. That is how basketball works.

range(shots$coordinate_x)   #> -46.75  46.75   <- length of the floor
range(shots$coordinate_y)   #> -25.00  25.00   <- width of the floor

So coordinate_x is not left-to-right, it is baseline-to-baseline. Both of my assumptions about the axes were wrong at once.

The fix folds both halves onto a single basket. Distance from the nearest baseline is 46.75 - abs(x), and the sideways position is y, sign-flipped on one half so left and right stay consistent after the rotation.

Before — plots a butterfly
shots <- pbp |>
  filter(shooting_play) |>
  transmute(x = coordinate_x, y = coordinate_y)
After — folds both ends onto one hoop

What convinced me the fold was right was not the picture. It was that the rim landed at 5.25 feet from the baseline — exactly where a real NBA hoop is. I did not put that number in. It fell out of the arithmetic. That is the best feeling in this whole project: when the data agrees with the physical world about something you never told it.

The bug that assertion didn't catch

The geometry check passed and I felt good about myself for about a week. Then I looked at the finished chart and the three-point line was cold. Not slightly — the entire arc was returning about 0.88 points per shot, which would mean the league shot 29% from three. That is not a number that happens.

The problem was how I decided whether a shot was worth two or three:

value = ifelse(score_value >= 3, 3L, 2L)   # wrong

score_value is the points a play scored. On a miss it is zero. So every missed three fell through to the else and got labelled a two. Made threes counted as threes, missed threes counted as missed twos, and the three-point zones got dragged down by tens of thousands of misattributed misses.

It is a nasty bug because the code does not look wrong, the script does not error, and the chart still renders a perfectly plausible court. It is only wrong if you know what the number should be.

The column I wanted was sitting right there:

pbp |> count(points_attempted)
#>   points_attempted        n
#> 1                2   137557
#> 2                3    98557

Populated on every row, made or missed. So now it is just value = as.integer(points_attempted), with a second assertion guarding it — and this one tests against reality rather than against my own reading of the code:

# League-wide this is ~55% on twos and ~36% on threes every single year.
stopifnot(
  between(fg_pct_2, 0.48, 0.62),
  between(fg_pct_3, 0.31, 0.40)
)

The build now prints 2PT 55.0% (1.10 pps) | 3PT 35.7% (1.07 pps) on every run, and refuses to write the files if it ever doesn't.

What I took from this

My geometry assertion tested the thing I had just fixed. It did not test the thing I hadn't thought of yet. Now when I find a bug I try to write the check for the category of mistake, not the specific one.

Two more things that were hiding

My standings table came out with 33 teams. Three of them were STARS, STRIPES and WORLD — the All-Star weekend rosters, filed by ESPN as if they were franchises. I filter them out by games played rather than by name, since real teams play 82 and exhibition rosters play two or three.

Then San Antonio and New York had 83 games. That one is the NBA Cup final: ESPN tags it as regular season, but it does not count in the standings. I could have hardcoded the game id, and nearly did. Instead I pull the schedule, which labels it directly, so the fix still works next season when it is two different teams:

Hexagons, and why not squares

Once the coordinates were right I binned them. Square bins make a shot chart look like a spreadsheet and produce grid artifacts that are not in the data. Hexagons tile without those seams, and every neighbour sits the same distance away.

Converting a point to a hex is a coordinate transform plus rounding, and the rounding is the part that is easy to get wrong. You cannot round the two axial coordinates independently, because the axes are not perpendicular and that misassigns points near the hex edges. You convert to cube coordinates, round all three, then repair whichever moved furthest.

219,607 shots go in. 243 hexagons come out, each carrying its attempt count, field goal percentage and points per shot:

Where the shots go, and whether they pay —
Hex area is attempts, colour is points per shot. The league view compares each bin to the league average of 1.08; a team view compares it to the league's own number from that same patch of floor. Hover any hex.

Why it all gets written to disk

Halfway through building this, a hoopR call failed on me with a 503 from GitHub — mid-pull, no warning. If that had happened while I was showing the app to someone, the demo would simply have been over.

So the script caches the raw download locally, every run after the first is completely offline, and the app never calls an API at all. It reads five JSON files totalling under 400 KB. The slow, fragile part happens once, on my machine, where a failure costs me thirty seconds instead of a meeting.

Every field is documented in DATA.md, and the full script is in the repo.

← All writing Next: AI as a build partner →