无需追踪实现动态3D场景的语义分割与快速编辑
TRASE: Tracking-free 4D Segmentation and Editing
- 基于弱监督学习构建4D语义特征场,利用SAM掩码引导对比学习
- 在五个动态数据集上实现跨视角的领先分割性能
- 支持对象移除、组合和风格迁移等实时交互编辑
理解动态3D场景对扩展现实(XR)和自动驾驶至关重要。将语义信息融入3D重建可实现完整场景表征,推动沉浸式与交互式应用的发展。为此,我们提出TRASE,一种新型无追踪4D分割方法,用于动态场景理解。TRASE以弱监督方式学习4D分割特征场,采用由SAM掩码引导的软挖掘对比学习目标。所得特征空间语义连贯且类间分离良好,最终通过无监督聚类获得物体级分割。该方法可直接操作场景中的高斯分布,实现快速编辑,如对象移除、组合和风格迁移。我们在五个动态基准上评估TRASE,证明其在未见视角下达到最优分割性能,并在多种交互编辑任务中表现优异。
原文摘要 · Abstract (English)
Understanding dynamic 3D scenes is crucial for extended reality (XR) and autonomous driving. Incorporating semantic information into 3D reconstruction enables holistic scene representations, unlocking immersive and interactive applications. To this end, we introduce TRASE, a novel tracking-free 4D segmentation method for dynamic scene understanding. TRASE learns a 4D segmentation feature field in a weakly-supervised manner, leveraging a soft-mined contrastive learning objective guided by SAM masks. The resulting feature space is semantically coherent and well-separated, and final object-level segmentation is obtained via unsupervised clustering. This enables fast editing, such as object removal, composition, and style transfer, by directly manipulating the scene's Gaussians. We evaluate TRASE on five dynamic benchmarks, demonstrating state-of-the-art segmentation performance from unseen viewpoints and its effectiveness across various interactive editing tasks. Our project page is available at: https://yunjinli.github.io/project-sadg/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。