← Back to the portfolio

Source & data pipeline

Colophon Everything here is reproducible

Every number on this site comes from a script in this repository. Nothing is sample data, nothing is hardcoded, and nothing is fetched from an API while you are looking at it.

That last part is deliberate. Halfway through building Hoop Vision, a hoopR call died with a 503 from GitHub in the middle of a pull. If that had been wired into a page instead of a build script, the failure would have happened in front of whoever I was showing it to. So the slow, fragile work happens once on my machine, the output gets committed, and every page here reads a static file.

The two pipelines

ScriptSourceWritesRows in
build_datasets.R hoopR / ESPN 5 JSON + data.js 642,472
build_nfl_qb.R nflfastR nfl.js 48,771

691,243 rows of play-by-play in; about 460 KB of JSON out. Both scripts cache their raw download to a gitignored .cache/, so the first run needs the network and every run after it does not.

Regenerating everything

Both scripts refuse to write anything if their assertions fail. Those assertions check the output against numbers from the real world rather than against my own reading of the code — league three-point percentage has to land between 31% and 40%, every NBA team has to have exactly 82 games, and the folded shot coordinates have to put the rim 5.25 feet from the baseline.

What's in the repo

.
├── index.html               # this portfolio
├── source.html              # you are here
├── blog/                    # four write-ups, each with inline charts
├── hoopvision/              # the Hoop Vision web app (static, no server)
│   ├── index.html
│   ├── app.js               # tabs, command palette, view wiring
│   └── app.css
├── assets/
│   ├── brand.css            # shared design tokens
│   ├── charts.js            # every SVG chart, hand-built, no library
│   ├── nflcharts.js         # the two NFL charts
│   ├── hero.js              # homepage dashboard
│   ├── code.js              # the view-source panels on this site
│   ├── data.js              # NBA bundle   (generated)
│   ├── nfl.js               # NFL bundle   (generated)
│   ├── fonts/               # self-hosted, so no CDN request
│   └── vendor/prism.js      # syntax highlighting, vendored not CDN
├── scripts/
│   └── build_nfl_qb.R       # NFL pipeline
└── Hoop_Vision-main/        # the R / Shiny implementation
    ├── scripts/build_datasets.R   # NBA pipeline — the only networked file
    ├── R/                   # modules, data layer, scoring, theme
    ├── data/                # committed JSON
    ├── tests/testthat/      # 23 tests
    ├── README.md
    └── DATA.md              # every field, and every gotcha

Two implementations of Hoop Vision

Worth being clear about this, because it looks like duplication and isn't quite.

The web app is what is deployed here. It is static HTML, CSS and hand-written SVG — no chart library and no server — because GitHub Pages only serves files. That is what makes it embeddable in the portfolio page as a same-origin iframe.

The R/Shiny app in Hoop_Vision-main/ is the original implementation, built against the spec docs in that folder using shinyMobile, echarts4r and reactable. It runs locally with shiny::runApp() and cannot be hosted on GitHub Pages, because Pages cannot run R.

Both read the same committed JSON from the same build script, so they cannot disagree about a number. The R version is the one with the test suite.

Charts

There is no charting library on this site. Every chart is SVG or canvas written by hand in assets/charts.js, assets/nflcharts.js and assets/hero.js. The categorical colour palette was checked against the dark surface for colour-vision separation and contrast rather than picked by eye.

Data attribution

NBA data comes from hoopR (SportsDataverse), which sources from ESPN. NFL data comes from nflfastR. Both publish prebuilt season releases, and this project reads those rather than scraping the underlying services. Please do the same.

← Back to the portfolio View the repo on GitHub →