How I Use AI as a Build Partner
Without Letting It Do the Thinking
I used an AI assistant for a lot of this project and I am not going to pretend otherwise. What I want to write down is the difference between the way I used it in week one, which made me worse, and the way I use it now, which makes me faster.
Week one: I was a copy-paste machine
My first approach was to describe what I wanted, get code back, paste it in, and see if it ran. If it ran, I moved on. If it errored, I pasted the error back and got a new version.
This felt incredibly productive. I had a working script by the end of the
first evening. It also meant that when someone asked me why I was using
load_nba_pbp() instead of one of the other loaders, I had no
answer, because I had never made that decision. Something made it for me and
I agreed by pasting.
The thing I eventually noticed is that "does it run" and "is it right" are almost unrelated questions in data work. A script that computes the wrong average runs perfectly. That is the whole problem.
The loop I use now
Four steps, and the third is the one that matters.
1. Prompt with the constraints I actually care about
Vague requests get plausible answers. Specific requests get checkable ones. So instead of "write me a function to bin shots into hexagons" I ask for the thing plus the properties it has to have — pure function, no reliance on row order, has to handle a point sitting exactly on a hex boundary. Naming what "correct" means up front is most of the work, and it is work only I can do, because I am the one who knows what the chart is for.
2. Read every line and make it explain itself
My rule is that I do not keep a line I cannot explain to someone else. When hex binning came back with cube-coordinate rounding, I genuinely did not understand why you could not just round the two axial coordinates. So I asked for the reason, then went and read the Red Blob Games hex guide to check the reason was real.
It was — rounding each axis on its own misassigns points near the corners of a hex, because the two axes are not perpendicular. Now I can explain it, so the code stays.
3. Try to break it
This is the step I skipped for the first two weeks and the only one that actually catches things.
The fantasy scoring function is a good example. It looked right. The weights matched the spec. So I hand-computed a line I could verify on paper and wrote it as an assertion instead of eyeballing it:
That one passed. Good. But writing it got me in the habit, and the habit is what caught the real bug later.
4. Re-prompt with what I learned, not with the error
Pasting a stack trace back gets you a patch. Explaining what you now understand about the problem gets you a fix. Those produce different code.
The afternoon I lost
Here is the one that actually taught me something.
I needed to know whether each shot was worth two points or three. I asked, and got back a line that looked completely reasonable:
value = ifelse(score_value >= 3, 3L, 2L)
I read it. I understood it. It made sense to me: if a play was worth three or more points, it is a three. I kept it. It ran fine. Every chart rendered.
Two weeks later I noticed the three-point line was cold on my shot chart — the whole arc sitting around 0.88 points per shot, which implies the league shoots 29% from deep. I know enough basketball to know that is absurd. The real number is around 36%.
score_value is the points a play scored. It is zero on
a miss. So every missed three fell through to the else branch
and was counted as a missed two.
What gets me about this is that it was not a hallucination. There was no made-up function, no wrong package, nothing an error message would ever catch. It was a reasonable reading of an ambiguously named column, and I reviewed it and agreed. Both of us were confidently wrong in exactly the same way.
Reviewing code for whether it makes sense is not the same as checking it against reality. I had been doing the first thing and telling myself it was the second.
The fix took ninety seconds — there is a points_attempted
column populated on every row, made or missed. Finding it took most of an
afternoon. And the real fix was not the line of code, it was adding a check
that compares the output against a number I know from outside the dataset:
# These land near 55% and 36% every season. If shot labelling
# breaks again, the build stops instead of drawing a wrong chart.
stopifnot(
between(fg_pct_2, 0.48, 0.62),
between(fg_pct_3, 0.31, 0.40)
)
It happened again on the NFL side, in a slightly different costume. My
quarterback completion percentages came out about four points too low,
because play_type == "pass" includes sacks and a sack is not a
pass attempt. Same shape of mistake: reasonable-looking code, no error, wrong
number, caught only because I knew roughly what the answer should be.
What I think it's actually good at
Being honest about both directions here.
It is genuinely excellent at things where I can verify the answer
immediately. Reshaping a data frame, remembering dplyr syntax I
half-know, explaining what an error means, writing the boring forty lines of
SVG axis code after I have decided what the axes should be. In those cases
the feedback loop is instant — either the shape is right or it is not.
It is most dangerous exactly where I am least able to check: domain assumptions inside a dataset I do not know well yet. That is where a wrong answer looks identical to a right one and survives review, because my review is only as good as my understanding.
The pattern is basically: the less I know about something, the less I should trust myself to review it, and the more I should go find an outside number to check against. Which is inconvenient, because those are also the moments when help is most tempting.
What I'd tell someone starting
Decide what correct means before you ask for anything. If you cannot describe how you would know the output was wrong, you are not ready to accept the output.
Check against the world, not against the code. My geometry assertion was right because the hoop landed at 5.25 feet — a real measurement I could look up. My shot values were wrong for two weeks because I never compared them to the actual league three-point percentage, which I could also have looked up in ten seconds.
And keep a list of everything you did not understand and looked up later.
Mine has maybe fifteen things on it — cube coordinates, why true shooting
uses 0.44, what renv is actually for, the difference between a
dropback and an attempt. That list is the part of this project I will still
have in five years.