EP 009

Short Stuff

Short Stuff: Evals

The least glamorous idea in applied AI and the one that separates systems that ship from systems that demo. What an evaluation set is, why twenty examples beats zero, and how to build one this afternoon.

0:0019 min

Audio for this episode is not attached yet — add an `audioUrl` to the episode file to enable playback

The least glamorous idea in applied AI and the one that separates systems that ship from systems that demo. What an evaluation set is, why twenty examples beats zero, and how to build one this afternoon.

In this episode

What an eval actually is

A set of inputs with known-good outputs. That is the whole idea, and almost nobody has one.

Twenty is enough to start

Why a small, honest, hand-labelled set beats a large synthetic one every time.

Building yours today

Pull twenty real cases from last month, write down what a good answer looks like, and run them before every change.

If you cannot measure whether a change made it better, you are not iterating. You are redecorating.

Mentioned


Beyond the Prompt publishes a new episode most weeks. Subscribe wherever you listen, or read the written edition first — both cover the same ground, at different speeds.

Keep listening

More episodes.

Behind the Work

Inside the Editorial Stack

We open our own machinery. The agents that draft, the editors that cut, the fact-checking pass that catches what the drafts invent, and the places where a human still has to sit down and decide.

Jul 15, 2026·53 min
Listen
Share Inside the Editorial Stack
Ideas

Learning Versus Laundering

There is a version of using AI that makes you sharper and a version that quietly hollows you out. They look identical from the outside and produce nearly identical documents. The difference is what you can do next week.

Jul 1, 2026·49 min
Listen
Share Learning Versus Laundering
Systems/with Daniel Okafor

Multi-Agent Systems, and When Not to Build One

Splitting a hard problem across specialised agents is genuinely powerful and usually premature. Daniel Okafor on coordination cost, the failure modes nobody warns you about, and the far simpler thing to try first.

Jun 24, 2026·56 min
Listen
Share Multi-Agent Systems, and When Not to Build One

The dispatch

One letter a week. Nothing else.

New essays, new episodes, and the occasional note about something we got wrong. No launch announcements, no course upsells, no breathless takes on last night's model release.

Unsubscribe in one click · We never sell the list