提出BVE框架与自建数据集,实现高效精准的3D文本编辑。
Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data

- 基于自构建数据集,用轻量模块注入文本语义
- 无需重训练即可实现高质量3D资产生成
- 免标注掩码策略保障未修改区域一致性
3D编辑指对3D资产进行局部或全局修改的能力。有效编辑需在提示指导下进行语义一致的局部调整,同时保持未修改区域的视觉一致性。现有方法存在明显局限:多视图编辑在3D投影回退时损失显著,体素编辑则受限于可修改区域和修改尺度。此外,缺乏足够大规模的训练与评估数据集仍是关键挑战。为此,我们提出超越体素的3D编辑框架(BVE),并构建了一个专为3D编辑设计的大规模自建数据集。基于该数据集,模型在基础图像到3D生成架构上引入轻量可训练模块,实现高效文本语义注入,无需昂贵的全模型重训练。同时,提出无标注3D掩码策略,以保持未修改区域的局部不变性。大量实验表明,BVE在生成高质量、文本对齐的3D资产方面表现优异,同时忠实保留原始输入的视觉特征。
原文摘要 · Abstract (English)
3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by performing localized changes according to prompts, while also preserving local invariance so that unchanged regions remain consistent with the original. However, existing approaches have significant limitations: multi-view editing methods incur losses when projecting back to 3D, while voxel-based editing is constrained in both the regions that can be modified and the scale of modifications. Moreover, the lack of sufficiently large editing datasets for training and evaluation remains a challenge. To address these challenges, we propose a Beyond Voxel 3D Editing (BVE) framework with a self-constructed large-scale dataset specifically tailored for 3D editing. Building upon this dataset, our model enhances a foundational image-to-3D generative architecture with lightweight, trainable modules, enabling efficient injection of textual semantics without the need for expensive full-model retraining. Furthermore, we introduce an annotation-free 3D masking strategy to preserve local invariance, maintaining the integrity of unchanged regions during editing. Extensive experiments demonstrate that BVE achieves superior performance in generating high-quality, text-aligned 3D assets, while faithfully retaining the visual characteristics of the original input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。