The Inside View cover image

5. Charlie Snell on DALL-E and CLIP

The Inside View

00:00

Dolly

This is how Dolly generates images. You input the text and then this transformer predicts a sequence of discrete latents. Then you pack that to the decoder to generate the actual image. So basically this is like effectively out of the box, the same thing as like a language model like GPT three or something.

Transcript
Play full episode

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!
App store bannerPlay store banner
Get the app