Text-to-image
Producing an image from a description. The approach that generalised the task treats text and image as one autoregressive stream of tokens, rather
Official source: Zero-Shot Text-to-Image Generation →
than adding auxiliary losses or side information — object part labels, segmentation masks — supplied during training. With enough data and scale it matched domain-specific models while being evaluated zero-shot, that is, without having been trained on the benchmark it was measured against.
Sources 1 source on record · Automated review 3 angles
Sources
What this page is based on — every source is verified and links out so you can check it yourself.
- Verified / official
- Provisional
- Out of date