实现多物体9自由度姿态精准控制,生成图像更灵活真实。
SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- 用新编码方式CNOCS地图表示9维姿态,提升训练稳定性。
- 在自建数据集ObjectPose9D上达到更高控制精度和图像质量。
- 支持个性化定制控制,适合需要精细布局的场景生成任务。
可控图像生成近年来备受关注,可实现对身份、风格等视觉内容的操控。然而,同时精确控制多个物体的9自由度姿态(位置、大小、方向)仍是难题。现有方法常受限于控制能力不足与画质下降。为此,我们提出SceneDesigner,一种高精度、灵活的多物体9-DoF姿态操控方法。该方法在预训练模型基础上引入分支网络,并使用新型表示CNOCS图,从相机视角编码9维姿态信息,具备强几何解释性,有助于高效稳定训练。为支持训练,我们构建了新数据集ObjectPose9D,整合多源图像并标注9维姿态。针对低频姿态性能下降问题,采用两阶段训练策略,第二阶段通过强化学习在重平衡数据上进行微调。推理时,提出解耦物体采样技术,缓解复杂场景中物体生成不足与概念混淆问题。结合用户个性化权重,支持参考主体的定制化姿态控制。大量定性和定量实验表明,SceneDesigner显著优于现有方法,在可控性与画质上均表现更优。代码已开源。
原文摘要 · Abstract (English)
Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simultaneous control over the 9D poses (location, size, and orientation) of multiple objects remains an open challenge. Despite recent progress, existing methods often suffer from limited controllability and degraded quality, falling short of comprehensive multi-object 9D pose control. To address these limitations, we propose SceneDesigner, a method for accurate and flexible multi-object 9-DoF pose manipulation. SceneDesigner incorporates a branched network to the pre-trained base model and leverages a new representation, CNOCS map, which encodes 9D pose information from the camera view. This representation exhibits strong geometric interpretation properties, leading to more efficient and stable training. To support training, we construct a new dataset, ObjectPose9D, which aggregates images from diverse sources along with 9D pose annotations. To further address data imbalance issues, particularly performance degradation on low-frequency poses, we introduce a two-stage training strategy with reinforcement learning, where the second stage fine-tunes the model using a reward-based objective on rebalanced data. At inference time, we propose Disentangled Object Sampling, a technique that mitigates insufficient object generation and concept confusion in complex multi-object scenes. Moreover, by integrating user-specific personalization weights, SceneDesigner enables customized pose control for reference subjects. Extensive qualitative and quantitative experiments demonstrate that SceneDesigner significantly outperforms existing approaches in both controllability and quality. Code is publicly available at https://github.com/FudanCVL/SceneDesigner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。