
Super Data Science: ML & AI Podcast with Jon Krohn 1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell
26 snips
Sep 18, 2026 Anton McGonnell, SambaNova’s VP of Product, explores why agentic AI is straining traditional inference hardware. He compares GPUs with SambaNova’s spatially mapped reconfigurable dataflow unit, designed for faster token generation and higher concurrency. They discuss the SN50’s economics, six-month payback target, air-cooled deployment, NeoClouds, sovereign infrastructure, and the rising cost of slow inference.
AI Snips
Chapters
Books
Transcript
Episode notes
Agentic AI Reshapes Inference Workloads
- Agentic AI has changed inference from short, simple prompts to large, reusable inputs and much larger KV caches during decoding.
- Hardware optimized for early generative AI workloads may struggle because agentic systems repeatedly process extensive context while generating outputs.
Inference Is Expanding Beyond Deployment
- Inference is strategically attractive because it is a larger, less concentrated market than training, where NVIDIA dominates major buyers.
- Post-training techniques such as reinforcement learning increasingly require inference compute, expanding demand beyond deployment alone.
Inference Economics Depend On Speed And Throughput
- Inference providers balance speed per user against throughput per chip, because faster responses reduce concurrency.
- SambaNova targets both sides of this Pareto trade-off, enabling providers to charge more for faster tokens while serving more users.




