The guests argue that most agent projects fail not on model capability but on evaluation: teams cannot tell whether a change made the agent better or worse, so they stop shipping changes.
The second half covers what has improved this year — replayable traces, cheaper models for grading, and sandboxed environments that make regression testing practical.