通过学习视角标记实现文本生成图像的精准相机控制。
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens

- 引入可学习的视角标记,实现全局场景理解下的相机控制。
- 在新数据集上微调模型,达到当前最佳精度且保持图像质量。
- 视角标记能泛化到未见物体类别,避免过拟合外观特征。
当前文本生成图像模型仅靠自然语言难以实现精确的相机控制。本文提出一种框架,通过学习参数化视角标记,在文本生成图像中实现精确的相机控制并具备全局场景理解能力。我们在一个精心构建的数据集上微调生成模型,该数据集结合了3D渲染图像以提供几何监督,以及逼真增强图像以丰富外观和背景多样性。定性和定量实验表明,本方法在保持图像质量和提示忠实度的同时,实现了最先进的准确率。与先前方法过度依赖特定物体外观关联不同,我们的视角标记学习的是可分离的几何表示,能够泛化到未见过的物体类别。研究证明,文本-视觉潜在空间可显式嵌入3D相机结构,为生成几何感知提示提供了新路径。
原文摘要 · Abstract (English)
Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text-to-image generation by learning parametric camera tokens. We fine-tune image generation models for viewpoint-conditioned text-to-image generation on a curated dataset that combines 3D-rendered images for geometric supervision and photorealistic augmentations for appearance and background diversity. Qualitative and quantitative experiments demonstrate that our method achieves state-of-the-art accuracy while preserving image quality and prompt fidelity. Unlike prior methods that overfit to object-specific appearance correlations, our viewpoint tokens learn factorized geometric representations that transfer to unseen object categories. Our work shows that text-vision latent spaces can be endowed with explicit 3D camera structure, offering a pathway toward geometrically-aware prompts for text-to-image generation. Project page: https://randdl.github.io/viewtoken_control/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。