arXiv:2511.13713cs.CV2025-11AAAI被引 3

让真实图像中的物体支持多轮3D操作,像在3D引擎里一样自然。

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine

  • 用序列化的3D变换建模编辑过程,避免逐帧重建。
  • 在多轮操作中保持光影与场景一致性,效果优于现有方法。
  • 适合需要精细3D编辑的设计师或工业应用者。

近期文本生成图像(T2I)扩散模型虽提升了语义图像编辑能力,但多数方法难以实现3D感知的物体操作。本文提出FFSE,一种3D感知的自回归框架,可在真实图像上直观、物理一致地编辑物体。不同于以往在图像空间操作或依赖耗时且易错的3D重建方法,FFSE将编辑建模为一系列学习到的3D变换,支持任意位移、缩放和旋转操作,并保留真实背景效果(如阴影、反射),同时在多轮编辑中维持全局场景一致性。为支持多轮3D感知编辑训练,我们构建了3DObjectEditor,一个融合模拟编辑序列的混合数据集,覆盖多样物体与场景。大量实验表明,该方法在单轮与多轮3D感知编辑任务中均显著优于现有方法。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive framework designed to enable intuitive, physically-consistent object editing directly on real-world images. Unlike previous approaches that either operate in image space or require slow and error-prone 3D reconstruction, FFSE models editing as a sequence of learned 3D transformations, allowing users to perform arbitrary manipulations, such as translation, scaling, and rotation, while preserving realistic background effects (e.g., shadows, reflections) and maintaining global scene consistency across multiple editing rounds. To support learning of multi-round 3D-aware object manipulation, we introduce 3DObjectEditor, a hybrid dataset constructed from simulated editing sequences across diverse objects and scenes, enabling effective training under multi-round and dynamic conditions. Extensive experiments show that the proposed FFSE significantly outperforms existing methods in both single-round and multi-round 3D-aware editing scenarios.

3D编辑扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。