Abstract
Despite recent significant progress on generative models, context-rich text-to-image synthesis depicting multiple complex objects is still non-trivial. The main challenges lie in the ambiguous semantic of a complex description and the intricate scene of an image with various objects, different positional relationship and diverse appearances. To address these challenges, we propose R-GAN, which can generate reasonable images according to the given text in a human-like way. Specifically, just like humans will first find and settle the essential elements to create a simple sketch, we first capture a monolithic-structural text representation by building a scene graph to find the essential semantic elements. Then, based on this representation, we design a bounding box generator to estimate the layout with position and size of target objects, and a following shape generator, which draws a fine-detailed shape for each object. Different from previous work only generating coarse shapes blindly, we introduce a coarse-to-fine shape generator based on a shape knowledge base. At last, to finish the final image synthesis, we propose a multi-modal geometry-aware spatially-adaptive generator conditioned on the monolithic-structural text representation and the geometry-aware map of the shapes. Extensive experiments on the real-world dataset MSCOCO show the superiority of our method in terms of both quantitative and qualitative metrics.
| Original language | English |
|---|---|
| Title of host publication | MM '21 |
| Subtitle of host publication | proceedings of the 29th ACM International Conference on Multimedia |
| Place of Publication | New York |
| Publisher | Association for Computing Machinery |
| Pages | 2085-2093 |
| Number of pages | 9 |
| ISBN (Electronic) | 9781450386517 |
| DOIs | |
| Publication status | Published - 2021 |
| Externally published | Yes |
| Event | 29th ACM International Conference on Multimedia, MM 2021 - Virtual, Online, China Duration: 20 Oct 2021 → 24 Oct 2021 |
Conference
| Conference | 29th ACM International Conference on Multimedia, MM 2021 |
|---|---|
| Country/Territory | China |
| City | Virtual, Online |
| Period | 20/10/21 → 24/10/21 |
Keywords
- Text-to-image Synthesis
- Generative Adversarial Networks
Fingerprint
Dive into the research topics of 'R-GAN: exploring human-like way for reasonable text-to-image synthesis via generative adversarial networks'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver