The Inside View cover image

5. Charlie Snell on DALL-E and CLIP

The Inside View

CHAPTER

Dolly

This is how Dolly generates images. You input the text and then this transformer predicts a sequence of discrete latents. Then you pack that to the decoder to generate the actual image. So basically this is like effectively out of the box, the same thing as like a language model like GPT three or something.

00:00
Transcript
Play full episode

Remember Everything You Learn from Podcasts

Save insights instantly, chat with episodes, and build lasting knowledge - all powered by AI.
App store bannerPlay store banner