33min chapter

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis cover image

Teaching AI to See: A Technical Deep-Dive on Vision Language Models with Will Hardman of Veratai

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

CHAPTER

Exploring Vision Language Models: Performance and Potential

This chapter explores the performance of vision language models (VLMs) in comparison to text-only benchmarks, revealing a general decline in capabilities with notable exceptions like the LAMA 3 series. It discusses the significance of quality datasets in fine-tuning and the impact of incorporating multimodal data on improving reasoning skills and mathematical problem-solving. The chapter also delves into the nuances of model integration, prompting strategies, and the future of AI architectures in harnessing diverse information modalities.

00:00

Get the Snipd
podcast app

Unlock the knowledge in podcasts with the podcast player of the future.
App store bannerPlay store banner

AI-powered
podcast player

Listen to all your favourite podcasts with AI-powered features

Discover
highlights

Listen to the best highlights from the podcasts you love and dive into the full episode

Save any
moment

Hear something you like? Tap your headphones to save it with AI-generated key takeaways

Share
& Export

Send highlights to Twitter, WhatsApp or export them to Notion, Readwise & more

AI-powered
podcast player

Listen to all your favourite podcasts with AI-powered features

Discover
highlights

Listen to the best highlights from the podcasts you love and dive into the full episode