Agents Keep Grading Their Own Homework (And Failing)
Episode 4 · 2026-06-14 · 4 min
Two weeks of AI agent news distilled into one conversation. Shrimp and Carl work through the reliability gap in production agent stacks: why self-verification is theater, why retries are quietly the dominant failure mode, why multi-agent debate can make answers worse, and why the interesting engineering has moved below the model layer entirely.