Dolly

This is how Dolly generates images. You input the text and then this transformer predicts a sequence of discrete latents. Then you pack that to the decoder to generate the actual image. So basically this is like effectively out of the box, the same thing as like a language model like GPT three or something.

Play episode from 52:33

Transcript

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!

Get the app