无需训练即可精准编辑3D模型,保持未修改区域一致性和整体连贯性。
VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- 在3D隐空间直接编辑,利用反演轨迹保留上下文特征
- 在保留区域中替换去噪特征,提升未编辑部分一致性
- 构建了人工标注的Edit3D-Bench数据集,验证方法优越性
3D局部区域编辑对游戏产业和机器人交互至关重要。现有方法通常在多视图渲染图像上编辑后再重建3D模型,难以精确保持未编辑区域的一致性与整体连贯性。受结构化3D生成模型启发,我们提出VoxHammer——一种无需训练的新方法,在3D隐空间实现精确且连贯的编辑。给定3D模型后,VoxHammer首先预测其反演轨迹,获取各时间步的隐变量与键值令牌。在去噪与编辑阶段,将保留区域的去噪特征替换为对应的反演隐变量和缓存键值令牌。通过保留这些上下文特征,确保未编辑区域的一致重建与编辑部分的自然融合。为评估保留区域的一致性,我们构建了Edit3D-Bench,一个包含数百个样本的人工标注数据集,每个样本均带有精细标注的3D编辑区域。实验表明,VoxHammer在未编辑区域一致性与整体质量方面显著优于现有方法。该方法有望用于合成高质量编辑配对数据,为上下文感知3D生成奠定数据基础。
原文摘要 · Abstract (English)
3D local editing of specified regions is crucial for game industry and robot interaction. Recent methods typically edit rendered multi-view images and then reconstruct 3D models, but they face challenges in precisely preserving unedited regions and overall coherence. Inspired by structured 3D generative models, we propose VoxHammer, a novel training-free approach that performs precise and coherent editing in 3D latent space. Given a 3D model, VoxHammer first predicts its inversion trajectory and obtains its inverted latents and key-value tokens at each timestep. Subsequently, in the denoising and editing phase, we replace the denoising features of preserved regions with the corresponding inverted latents and cached key-value tokens. By retaining these contextual features, this approach ensures consistent reconstruction of preserved areas and coherent integration of edited parts. To evaluate the consistency of preserved regions, we constructed Edit3D-Bench, a human-annotated dataset comprising hundreds of samples, each with carefully labeled 3D editing regions. Experiments demonstrate that VoxHammer significantly outperforms existing methods in terms of both 3D consistency of preserved regions and overall quality. Our method holds promise for synthesizing high-quality edited paired data, thereby laying the data foundation for in-context 3D generation. See our project page at https://huanngzh.github.io/VoxHammer-Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。