Vidrial: A CUDA Framework for Fast Kernels

Jacob explains building Vidrial to write efficient, testable CUDA kernels and how it outperforms Triton/FlashAttention in shapes.

Play episode from 35:52

Transcript

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!