← All posts Podcast

Agents that actually ship

Sep 11, 20265 min read

Forty-five minutes on eval harnesses, why most agent demos die in production, and the tooling gap that is finally closing.

The guests argue that most agent projects fail not on model capability but on evaluation: teams cannot tell whether a change made the agent better or worse, so they stop shipping changes.

The second half covers what has improved this year — replayable traces, cheaper models for grading, and sandboxed environments that make regression testing practical.

Key points

  • Evaluation, not model quality, is the usual reason agents stall
  • Replayable traces and cheap graders make regression tests practical
  • Agent evals is a rising topic across sources this week