arXiv:2602.05676cs.CVcs.GR2026-02International Conf…被引 3

用图像提示实现高效3D编辑,保持结构一致且可扩展。

ShapeUP: Scalable Image-Conditioned 3D Editing

  • 通过3D扩散变换器,将编辑图像映射到3D潜在空间。
  • 在身份保留和编辑保真度上超越现有训练与免训练基线。
  • 支持细粒度控制,无需掩码即可定位修改区域。

近期3D基础模型的进步使得高质量资产生成成为可能,但精确的3D操作仍具挑战。现有3D编辑框架常面临视觉可控性、几何一致性与可扩展性之间的权衡:基于优化的方法过于缓慢,多视角2D传播技术存在视觉漂移,而免训练的潜在空间操作受限于冻结先验,难以受益于规模扩展。本文提出ShapeUP,一种可扩展的图像条件3D编辑框架,将编辑建模为原生3D表示中的监督潜在到潜在映射。该方法利用预训练3D基础模型的强大生成先验,并通过监督训练适配其编辑能力。具体地,ShapeUP在由源3D形状、编辑后的2D图像及对应编辑后的3D形状组成的三元组上训练,采用3D扩散变压器(DiT)学习直接映射。该图像作为提示的方法实现了对局部与全局编辑的细粒度视觉控制,实现隐式、无掩码的定位,同时严格保持原始资产的结构一致性。大量评估表明,ShapeUP在身份保留和编辑保真度方面持续优于当前的训练与免训练基线,为原生3D内容创作提供了一种鲁棒且可扩展的新范式。

原文摘要 · Abstract (English)

Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant challenge. Existing 3D editing frameworks often face a difficult trade-off between visual controllability, geometric consistency, and scalability. Specifically, optimization-based methods are prohibitively slow, multi-view 2D propagation techniques suffer from visual drift, and training-free latent manipulation methods are inherently bound by frozen priors and cannot directly benefit from scaling. In this work, we present ShapeUP, a scalable, image-conditioned 3D editing framework that formulates editing as a supervised latent-to-latent translation within a native 3D representation. This formulation allows ShapeUP to build on a pretrained 3D foundation model, leveraging its strong generative prior while adapting it to editing through supervised training. In practice, ShapeUP is trained on triplets consisting of a source 3D shape, an edited 2D image, and the corresponding edited 3D shape, and learns a direct mapping using a 3D Diffusion Transformer (DiT). This image-as-prompt approach enables fine-grained visual control over both local and global edits and achieves implicit, mask-free localization, while maintaining strict structural consistency with the original asset. Our extensive evaluations demonstrate that ShapeUP consistently outperforms current trained and training-free baselines in both identity preservation and edit fidelity, offering a robust and scalable paradigm for native 3D content creation.

3D编辑扩散模型图像引导可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。