
5. Charlie Snell on DALL-E and CLIP
The Inside View
Dolly
This is how Dolly generates images. You input the text and then this transformer predicts a sequence of discrete latents. Then you pack that to the decoder to generate the actual image. So basically this is like effectively out of the box, the same thing as like a language model like GPT three or something.
00:00
Transcript
Play full episode
Remember Everything You Learn from Podcasts
Save insights instantly, chat with episodes, and build lasting knowledge - all powered by AI.