让3D场景自动识别用户查询的子概念,兼顾语义与用户意图。
DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF
- 基于神经场,融合无监督分割与弱开集引导。
- 在开集与无监督场景下均达到顶尖性能。
- 适合需要灵活理解3D场景的机器人应用。
3D语义分割为机器人、自动驾驶等应用提供高层场景理解。传统方法仅适配特定任务目标(开集分割)或场景内容(无监督分割)。我们提出DiSCO-3D,首个解决3D开集子概念发现的通用方法,旨在实现既适应场景又响应用户查询的3D语义分割。该方法基于神经场表示,结合无监督分割与弱开集引导。实验表明,DiSCO-3D在开集子概念发现任务中表现优异,并在开集与无监督分割的边缘案例中均取得当前最优结果。
原文摘要 · Abstract (English)
3D semantic segmentation provides high-level scene understanding for applications in robotics, autonomous systems, \textit{etc}. Traditional methods adapt exclusively to either task-specific goals (open-vocabulary segmentation) or scene content (unsupervised semantic segmentation). We propose DiSCO-3D, the first method addressing the broader problem of 3D Open-Vocabulary Sub-concepts Discovery, which aims to provide a 3D semantic segmentation that adapts to both the scene and user queries. We build DiSCO-3D on Neural Fields representations, combining unsupervised segmentation with weak open-vocabulary guidance. Our evaluations demonstrate that DiSCO-3D achieves effective performance in Open-Vocabulary Sub-concepts Discovery and exhibits state-of-the-art results in the edge cases of both open-vocabulary and unsupervised segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。