DraftKings scoring · rolling window · minimum 30 games and 18 minutes per game
⌕
# Weights live here as data, not buried inside a chart. Matches
# docs/FANTASY_SPEC.md, and they get written into meta.json so the app
# can show exactly which numbers produced a value on screen.
DK_WEIGHTS <- c(points = 1, three_point_field_goals_made = 0.5, rebounds = 1.25,
assists = 1.5, steals = 2, blocks = 2, turnovers = -0.5)
DK_BONUS <- c(double_double = 1.5, triple_double = 3)
score_draftkings <- function(df) {
cats <- df |> select(points, rebounds, assists, steals, blocks)
# a "double" is 10+ in any of the five main categories
n_doubles <- rowSums(cats >= 10, na.rm = TRUE)
base <- with(df,
points * DK_WEIGHTS[["points"]] +
three_point_field_goals_made * DK_WEIGHTS[["three_point_field_goals_made"]] +
rebounds * DK_WEIGHTS[["rebounds"]] + assists * DK_WEIGHTS[["assists"]] +
steals * DK_WEIGHTS[["steals"]] + blocks * DK_WEIGHTS[["blocks"]] +
turnovers * DK_WEIGHTS[["turnovers"]])
df |> mutate(fantasy_pts = base +
ifelse(n_doubles >= 2, DK_BONUS[["double_double"]], 0) +
ifelse(n_doubles >= 3, DK_BONUS[["triple_double"]], 0))
}
# Checked against a line I can do on paper rather than trusted:
# 30 pts / 10 reb / 10 ast / 2 stl = 30 + 12.5 + 15 + 4 = 61.5,
# plus 1.5 (double-double) + 3 (triple-double) = 66.0
stopifnot(abs(score_draftkings(tibble(points = 30, rebounds = 10, assists = 10,
steals = 2, blocks = 0, turnovers = 0,
three_point_field_goals_made = 0))$fantasy_pts - 66.0) < 1e-9)
Skill fingerprint
Percentile rank against the 284 qualified players — press ⌘K to find anyone
Six axes, all percentiles rather than raw numbers — 25 points and 8 assists
aren't on the same scale, so a radar of raw stats would only show which stat has bigger numbers.
Usage is a proxy (possessions ended per 36 minutes), not official USG%.
# Percentile rank within the qualified pool. Comparing raw per-36 numbers
# across categories is meaningless (25 pts and 8 ast aren't the same scale);
# percentiles are.
pct_rank <- function(x) round(100 * rank(x, ties.method = "average") / length(x))
players <- season_avg |>
mutate(
p_scoring = pct_rank(pts36),
p_rebound = pct_rank(reb36),
p_passing = pct_rank(ast36),
p_defense = pct_rank(stk36), # steals + blocks per 36
p_effic = pct_rank(ts_pct), # true shooting
p_usage = pct_rank(usage36) # possessions ended per 36 -- a proxy
) |>
arrange(desc(fppg)) |>
# rowwise() matters here. Without it, c(p_scoring, p_rebound, ...) glues
# every player's percentiles into one giant vector and each row gets a copy
# of the whole thing -- players.json was 1.4 MB before I caught this.
rowwise() |>
mutate(radar = list(c(p_scoring, p_rebound, p_passing,
p_defense, p_effic, p_usage))) |>
ungroup()
The scoring race, game by game
Cumulative points · top 8 scorers · x-axis is each player's own Nth game
Plotting each player's own game number instead of the calendar date matters:
teams don't play the same nights, and someone who misses three weeks would otherwise look like
they flatlined. A short line is a player who missed games, not one who stopped scoring.
/* The top eight finish within ~240 points of each other, which is about
40px of chart. Eight headshots stacked in 40px is a smudge, so the
avatars get pushed apart vertically and a leader line runs back to the
real end of each line. The dot is the truth; the avatar is the label. */
const GAP = 34;
const ordered = [...ends].sort((a, b) => a.ey - b.ey);
ordered.forEach((e, k) => {
e.ty = k === 0 ? e.ey : Math.max(e.ey, ordered[k - 1].ty + GAP);
});
const overflow = ordered[ordered.length - 1].ty - (M.t + ih);
if (overflow > 0) {
ordered.forEach((e) => { e.ty -= overflow; });
for (let k = ordered.length - 2; k >= 0; k--) {
ordered[k].ty = Math.min(ordered[k].ty, ordered[k + 1].ty - GAP);
}
}
Shot frequency & efficiency
219,607 field-goal attempts, binned into hexagons
Every hexagon is a patch of floor. Size is how often shots
get taken there. Colour is whether they pay off.
ESPN reports shot coordinates for the whole floor, so a team's attempts are
split across both baskets depending on which way they were shooting. My first version came out
as two mirrored blobs with a hole at halfcourt. Folding both ends onto one hoop puts the rim at
5.25 feet from the baseline — exactly where a real one is, which is how I knew the fix was right.
# THE GOTCHA. ESPN gives shot coordinates for the WHOLE floor:
# coordinate_x in [-46.75, 46.75] <- length of the court, baseline to baseline
# coordinate_y in [-25, 25] <- width of the court
#
# So a team's shots are split across both ends depending on which way they were
# shooting that half. My first hexbin looked like a butterfly -- two mirrored
# blobs with a dead zone at halfcourt.
HOOP_FROM_BASELINE <- 5.25
COURT_HALF_LEN <- 46.75
shots_raw <- pbp |>
filter(season_type == REG_SEASON, shooting_play,
!is.na(coordinate_x), !is.na(coordinate_y),
!grepl("Free Throw", type_text, fixed = TRUE)) |>
transmute(
team_id,
made = as.logical(scoring_play),
# Use points_attempted, NOT score_value. score_value is the points actually
# scored, so it's 0 on every miss -- which quietly labeled every missed
# three as a two and dragged 3PT zones down to 0.88 points per shot.
value = as.integer(points_attempted),
# fold both ends of the floor onto one half-court
across_ft = ifelse(coordinate_x > 0, coordinate_y, -coordinate_y),
up_ft = COURT_HALF_LEN - abs(coordinate_x)
) |>
mutate(dist = sqrt(across_ft^2 + (up_ft - HOOP_FROM_BASELINE)^2))
# Verify the fold worked: threes start at 22 ft in the corners, 23.75 above
# the break. If this fails the geometry is wrong and the build stops.
three_p05 <- quantile(shots_raw$dist[shots_raw$value == 3], 0.05)
stopifnot(three_p05 > 20, three_p05 < 25)
Final standings
Click any column to sort · the sparkline is cumulative win% across all 82 games
Getting to exactly 30 teams took two fixes. ESPN files the All-Star rosters
(Stars, Stripes, World) as if they were franchises, and it tags the NBA Cup final as a
regular-season game — which is why San Antonio and New York first showed up with 83.
# Two surprises in ESPN's "regular season" bucket that I only caught because
# my first standings table had 33 teams in it:
#
# * The All-Star weekend rosters (STARS, STRIPES, WORLD) are in here as if
# they were franchises. They play 2-3 "games". Real teams play 82, so I
# filter on games played rather than hardcoding a list of 30 abbreviations.
#
# * The NBA Cup Championship game is tagged season_type 2, but it does NOT
# count toward the standings. Instead of hardcoding that game's id, I pull
# the schedule and drop anything ESPN labels as the Cup final -- so this
# keeps working next season when it's two different teams.
nba_team_ids <- team_box |>
filter(season_type == REG_SEASON) |>
count(team_id, name = "games") |>
filter(games >= 70) |>
pull(team_id)
cup_final_ids <- schedule |>
filter(grepl("Cup Championship", notes_headline, ignore.case = TRUE)) |>
pull(game_id) |>
unique()
countable <- function(df) {
df |> filter(team_id %in% nba_team_ids, !game_id %in% cup_final_ids)
}