arXiv:2606.03314cs.CV2026-06

让3D场景编辑更灵活可控,支持文本驱动的大规模修改。

TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing

论文配图:TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing
图 1 · 摘自论文原文
  • 将2D语义特征映射到可裁剪的嵌入空间,渐进减少通道生成抽象表示。
  • 在大规模几何修改下,编辑效果优于现有方法,且能保留细节或抽象内容。
  • 适合需要高自由度场景编辑的开发者,尤其适用于机器人与自动驾驶仿真。

高保真语义3D场景表示对机器人、自动驾驶和仿真等应用至关重要。现有方法在可控编辑方面支持有限。本文提出TASE,通过将预训练2D语义特征投影至截断感知嵌入空间,实现灵活的3D场景编辑。该方法显式优化特征空间,使逐步减少通道数产生越来越抽象的语义表示,保留更多通道则维持精细细节。同时,引入尺度与平移等变损失提升多视角一致性。所获截断感知嵌入空间支持文本驱动编辑,可明确控制修改强度,允许比以往方法更大幅度的改动。此外,我们为编辑扩散模型设计了微调阶段,以缓解几何变化带来的伪影。实验表明,该方法在3D场景编辑任务中表现优异,尤其在涉及大范围几何修改时显著优于先前方法。

原文摘要 · Abstract (English)

High-fidelity semantic 3D scene representations are crucial for numerous applications, including robotics, autonomous driving, and simulation. Beyond this, the ability to edit such representations enables developers to adapt these applications more easily to specific target scenarios. Current approaches provide limited support for controllable editing. We introduce TASE, a method that projects pretrained 2D semantic features into a truncation-aware embedding space to enable flexible 3D scene editing. Our method explicitly optimizes a feature space in which progressively reducing feature channels yields increasingly abstract semantic representations, while retaining more channels preserves fine-grained detail. Additionally, we improve multi-view consistency of the features using a scale- and translation-equivariance loss. The resulting truncation-aware embedding space enables text-driven edits to 3D scenes, providing explicit control over how strongly edits adhere to the original scene content and allowing more substantial modifications than prior methods. Moreover, we propose a finetuning stage for the editing diffusion model to mitigate artifacts caused by geometric changes. Experimental results demonstrate competitive performance in 3D scene editing, substantially outperforming prior methods on edits involving large geometric modifications.

3D编辑语义嵌入扩散模型场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。