A reliability layer that decomposes an LLM response into atomic claims, checks them against retrieved evidence, and returns a confidence score with the reasoning attached.
Things I’ve built,
and things I’ve written.
Two kinds of work live here. Products I built or shipped, each with the prototype and the write-up on one page. And the writing, where I think out loud about product and AI.
BUILTPrototype and write-up, on the same page.
A music discovery product that ranks on what a song sounds like rather than on how many people already found it.
A working harness for evaluating AI product changes — fixed task sets, explicit graders, and run-over-run comparison.
More coming here shortly.
WRITTENWhere the thinking happens, before it becomes a product.
Nothing published here yet. The first pieces are being moved over, including an article on buttons and prompts written for Bootcamp, a product and UX publication with 80K+ followers.