用粗略框选区域实现高质量3D物体编辑,更贴近真实操作习惯。
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

- 基于粗略3D框和参考2D图进行区域感知编辑
- 在多个数据集上显著优于现有方法的视觉与量化指标
- 适合需要快速修改3D模型局部细节的设计师或开发者
3D物体的局部编辑仍是长期挑战。人类在交互时通常仅指定粗略感兴趣区域,而非精确边界。但以往方法依赖完整2D图像、精确3D掩码或冗余流程,存在差距。为此,我们提出EditVerse3D,一种新型3D编辑框架,可在粗略引导下实现高质量编辑。输入包括待编辑3D对象、目标区域的粗略3D边界框及描述期望修改的参考2D图像,输出一致且高保真的编辑结果。为支持此任务,我们引入区域感知自适应损失,强调难学习区域并平衡目标区与保留区间的优化目标。同时,通过缩放3D掩码训练和剔除不现实编辑对等数据增强提升模型鲁棒性与泛化能力。我们构建了一个基于部件信息的大规模3D编辑数据集。大量实验表明,EditVerse3D在视觉质量和定量性能上均优于现有3D编辑方法。
原文摘要 · Abstract (English)
Local editing of 3D objects remains a long-standing challenge. When interacting with 3D content, humans naturally tend to specify a coarse region of interest for modification rather than defining precise editing boundaries. However, previous methods rely on fully edited 2D images, precise 3D masks, or redundant pipelines, which present a gap. To bridge this gap, we propose EditVerse3D, a novel 3D editing framework that enables high-quality object editing under such coarse guidance. Our approach takes as input a 3D object to be edited, a coarse 3D bounding box indicating the target region, and a reference 2D image describing the desired modification. It produces a coherent, high-fidelity edited 3D object. To facilitate this editing, we introduce a novel region-aware adaptive loss that emphasizes hard-to-learn regions and balances the objective between target and preserved areas. Complementing our loss function, we enhance model robustness and generalization through targeted data augmentations, such as training with scaled 3D masks and filtering out unrealistic editing pairs. We construct a large-scale 3D editing dataset derived from parts information. Extensive experiments demonstrate that EditVerse3D achieves superior visual quality and quantitative performance compared to existing 3D editing approaches. Please visit our project page at https://editverse3d.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。