FlexGen通过文本与图像输入生成可控多视角图像,支持灵活修改外观和材质。
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
- 利用GPT-4V生成带3D空间关系的文本标注,增强多视角控制能力。
- 通过自适应双控制模块,实现文本与图像条件下的多视图一致性生成。
- 支持材质、金属度等属性调节,适合游戏与VR快速内容创作。
本文提出FlexGen,一种灵活的多视图生成框架,可基于单视图图像、文本提示或两者结合生成可控且一致的多视角图像。通过引入3D感知文本标注作为额外控制信号,利用GPT-4V分析四张正交视图组成的多视图图像,生成包含空间关系的3D-aware文本描述。将该控制信号与提出的自适应双控制模块融合,模型可生成与指定文本对应的一致多视图图像。该方法支持多种可控属性,用户可通过修改文本提示生成合理且未见过的部分,还可调控外观及材质属性(如金属感、粗糙度)。大量实验表明,本方法在多模态可控性方面显著优于现有扩散模型,对游戏开发、动画制作与虚拟现实等需快速生成3D内容的领域具有重要价值。
原文摘要 · Abstract (English)
In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware text annotations. We utilize the strong reasoning capabilities of GPT-4V to generate 3D-aware text annotations. By analyzing four orthogonal views of an object arranged as tiled multi-view images, GPT-4V can produce text annotations that include 3D-aware information with spatial relationship. By integrating the control signal with proposed adaptive dual-control module, our model can generate multi-view images that correspond to the specified text. FlexGen supports multiple controllable capabilities, allowing users to modify text prompts to generate reasonable and corresponding unseen parts. Additionally, users can influence attributes such as appearance and material properties, including metallic and roughness. Extensive experiments demonstrate that our approach offers enhanced multiple controllability, marking a significant advancement over existing multi-view diffusion models. This work has substantial implications for fields requiring rapid and flexible 3D content creation, including game development, animation, and virtual reality. Project page: https://xxu068.github.io/flexgen.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。