让文字生成图像时精准控制物体3D朝向,无需额外训练。
ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
- 用奖励引导采样+随机噪声注入,实现3D朝向控制。
- 在多个物体类别上超越现有方法,用户评分更高。
- 零样本设计,适合需快速调整图像朝向的场景。
我们提出ORIGEN,首个可在多种物体和多样化类别中实现文本到图像生成的零样本3D朝向定位方法。以往图像生成中的空间定位研究主要聚焦2D位置控制,缺乏对3D朝向的有效调控。为此,我们采用基于预训练判别模型的3D朝向估计,结合单步文本到图像生成流程,并引入奖励引导采样策略。不同于依赖梯度上升的优化方式,该方法通过朗之万动力学采样,在仅增加一行代码的前提下注入随机噪声,有效保持图像真实感。此外,我们提出基于奖励函数的自适应时间重缩放机制,加速收敛。实验表明,ORIGEN在定量指标和用户评测中均优于基于训练和测试时引导的方法。
原文摘要 · Abstract (English)
We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address this, we propose a reward-guided sampling approach using a pretrained discriminative model for 3D orientation estimation and a one-step text-to-image generative flow model. While gradient-ascent-based optimization is a natural choice for reward-based guidance, it struggles to maintain image realism. Instead, we adopt a sampling-based approach using Langevin dynamics, which extends gradient ascent by simply injecting random noise--requiring just a single additional line of code. Additionally, we introduce adaptive time rescaling based on the reward function to accelerate convergence. Our experiments show that ORIGEN outperforms both training-based and test-time guidance methods across quantitative metrics and user studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。