The AWS Developers Podcast

Why Your Agent Evaluations Will Fail You (and How to Fix Them Before Production)

8 snips
Jun 3, 2026
James Price-Farr, AI Engineering Team Lead at Xelix (builds production ML and agentic systems), and Paul Solomon, Head of AI Engineering at Xelix (scales AI/ML for enterprise accounts payable). They explain why evaluating tool calls beats just checking outputs. They cover three automation tiers, steering files vs prompts, orchestration pitfalls, Bedrock rollout lessons, and scaling automation from 10% up.
Ask episode
AI Snips
Chapters
Books
Transcript
Episode notes
INSIGHT

Agents Solve The Long Tail Of Email Cases

  • Agents excel at handling the long tail of variable, edge-case supplier emails that single LLM calls struggle with.
  • Xelix exposes per-customer tools and data to an agent so it can fetch DB records and documents dynamically.
ADVICE

Build Evaluation And Monitoring From Day One

  • Do build evaluation and monitoring from day one for agent workflows.
  • James Price-Farr logs which tools an agent calls and correlates tool-usage stats with email categories to detect failures early.
INSIGHT

Design Tasks To Be Self Validating

  • Designing tasks with self-validation reduces risk for touchless automation.
  • Xelix structures document-processing tasks so success is obvious and failures naturally block automation, aiding reliability.
Get the Snipd Podcast app to discover more snips from this episode
Get the app