about
I started in engineering, crossed into product, and ended up spending most of my time on agentic systems and the evals that keep them honest. That's the loudest part of my work right now, but not the only part — I've shipped across checkout, identity, and commerce at scale, and I care a lot about the unsexy stuff like measurement, trust, and what “working” actually means in production. Outside of work, I run half-marathons (slowly, happily), read more than is probably reasonable, and have a well-documented soft spot for elephants.
eval results · hold-out set: me
0.00
PRECISION ON "THIS FIX WILL TAKE 10 MINS"
1:0
BLOG DRAFT → PUBLISH RATIO
0
ACTIVE SIDE PROJECTS
0
BOOKS READ THIS YEAR
0.0
WEBSITE VERSION
work
download resumeProducts, platforms, and experiences I’ve helped build and scale.
appearances & accolades
Talks, panels, and industry honors from along the way.
projects & experiments
Sandboxes, weekend builds and side projects testing new tech.
publications
Peer-reviewed research on machine learning and large language models.
bookmarks
A curated collection of links worth revisiting.
Evals are the AI buzzword for 2026. This article is foundational.
Chip Huyen on the gap between demos and production. The most grounded take on evaluation I've found.
This article played a big role when I decided to switch to product management. Good framing for PM-eng collaboration.
PG at his most earnest. The advice about following your curiosity over prestige is the one I keep coming back to.
Lilian Weng's definitive survey of LLM agent architectures. Dense, precise, worth re-reading every 6 months.
Sutton's argument that general methods leveraging compute always win. Required reading before any architecture debate.