arXiv:2510.23956cs.CVcs.AI2025-10被引 2

让3D场景编辑像拼积木一样精准可控

Neural USD: An object-centric framework for iterative editing and control

  • 用分层结构表示物体与场景,支持逐对象控制
  • 可实现颜色、姿态等属性的独立修改而不影响整体
  • 适合需要反复迭代修改的3D内容创作人群

可控生成模型虽进展显著,但精确且迭代式的目标编辑仍存挑战。现有方法通过调整条件信号修改图像(如改变某物体颜色或背景),常引发全局意外变化。本文受计算机图形学中通用场景描述标准USD启发,提出「神经通用场景描述」(Neural USD)框架,以结构化分层方式表示场景与物体,兼容多种信号输入,减少模型依赖,实现对物体外观、几何和姿态的逐对象控制。通过微调策略,确保各控制信号相互解耦。我们验证了多个设计选项,证明该框架支持迭代与增量式工作流程。更多信息见:https://escontrela.me/neural_usd。

原文摘要 · Abstract (English)

Amazing progress has been made in controllable generative modeling, especially over the last few years. However, some challenges remain. One of them is precise and iterative object editing. In many of the current methods, trying to edit the generated image (for example, changing the color of a particular object in the scene or changing the background while keeping other elements unchanged) by changing the conditioning signals often leads to unintended global changes in the scene. In this work, we take the first steps to address the above challenges. Taking inspiration from the Universal Scene Descriptor (USD) standard developed in the computer graphics community, we introduce the "Neural Universal Scene Descriptor" or Neural USD. In this framework, we represent scenes and objects in a structured, hierarchical manner. This accommodates diverse signals, minimizes model-specific constraints, and enables per-object control over appearance, geometry, and pose. We further apply a fine-tuning approach which ensures that the above control signals are disentangled from one another. We evaluate several design considerations for our framework, demonstrating how Neural USD enables iterative and incremental workflows. More information at: https://escontrela.me/neural_usd .

3D生成可控编辑分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。