Super Data Science: ML & AI Podcast with Jon Krohn

1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell

26 snips
Sep 18, 2026
Anton McGonnell, SambaNova’s VP of Product, explores why agentic AI is straining traditional inference hardware. He compares GPUs with SambaNova’s spatially mapped reconfigurable dataflow unit, designed for faster token generation and higher concurrency. They discuss the SN50’s economics, six-month payback target, air-cooled deployment, NeoClouds, sovereign infrastructure, and the rising cost of slow inference.
Ask episode
AI Snips
Chapters
Books
Transcript
Episode notes
INSIGHT

Agentic AI Reshapes Inference Workloads

  • Agentic AI has changed inference from short, simple prompts to large, reusable inputs and much larger KV caches during decoding.
  • Hardware optimized for early generative AI workloads may struggle because agentic systems repeatedly process extensive context while generating outputs.
INSIGHT

Inference Is Expanding Beyond Deployment

  • Inference is strategically attractive because it is a larger, less concentrated market than training, where NVIDIA dominates major buyers.
  • Post-training techniques such as reinforcement learning increasingly require inference compute, expanding demand beyond deployment alone.
INSIGHT

Inference Economics Depend On Speed And Throughput

  • Inference providers balance speed per user against throughput per chip, because faster responses reduce concurrency.
  • SambaNova targets both sides of this Pareto trade-off, enabling providers to charge more for faster tokens while serving more users.
Get the Snipd Podcast app to discover more snips from this episode
Get the app