让图像物体操作更符合物理规律,实现精准空间调整。
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

- 用显式3D几何模拟提供视觉引导,结合2D-3D联合监督。
- 在真实世界数据集上,3D几何精度和操作一致性显著提升。
- 适合需要真实物理交互的生成模型研究者与开发者。
在图像编辑中实现物理准确的物体操作对交互式世界模型应用至关重要。然而现有视觉生成模型在精确空间操作上表现不佳,常出现物体缩放与位置错误。这一局限主要源于缺乏显式的3D几何与透视投影机制。为此,我们提出PhyEdit框架,通过显式几何模拟作为上下文3D感知的视觉引导,并结合2D-3D联合监督,有效提升物理准确性与操作一致性。为支持该方法并评估性能,我们构建了真实世界数据集RealManip-40K,包含成对图像与深度标注;并提出ManipEval基准,涵盖多维度指标以评估3D空间控制与几何一致性。大量实验表明,该方法在3D几何精度与操作一致性方面超越现有方法,包括强闭源模型。
原文摘要 · Abstract (English)
Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However, existing visual generative models often fail at precise spatial manipulation, resulting in incorrect scaling and positioning of objects. This limitation primarily stems from the lack of explicit mechanisms to incorporate 3D geometry and perspective projection. To achieve accurate manipulation, we develop PhyEdit, an image editing framework that leverages explicit geometric simulation as contextual 3D-aware visual guidance. By combining this plug-and-play 3D prior with joint 2D--3D supervision, our method effectively improves physical accuracy and manipulation consistency. To support this method and evaluate performance, we present a real-world dataset, RealManip-40K, for 3D-aware object manipulation featuring paired images and depth annotations. We also propose ManipEval, a benchmark with multi-dimensional metrics to evaluate 3D spatial control and geometric consistency. Extensive experiments show that our approach outperforms existing methods, including strong closed-source models, in both 3D geometric accuracy and manipulation consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。