构建开放词汇3D场景理解新基准,支持自然语言多样化描述
OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
- 引入同义词与细粒度描述,标签数达原数据集13倍
- 设计开集语义分割与物体检索任务,暴露现有方法缺陷
- 适合研究开放词汇3D理解的学者与工业界应用开发者
3D场景理解因开放词汇语言模型的应用而革新,实现了通过自然语言交互。然而当前评估仍局限于封闭类别语义的数据集,无法体现语言的丰富性。本文提出OpenLex3D,一个专门用于评估3D开放词汇场景表示的基准。该基准为Replica、ScanNet++和HM3D中的场景提供全新标注,通过引入同义物类别和更细致的描述,捕捉真实语言多样性,每场景标签数是原数据集的13倍。通过引入开集3D语义分割与物体检索任务,我们在OpenLex3D上评估多种现有3D开放词汇方法,揭示其失败案例与改进方向。实验揭示了特征精度、分割效果及下游能力的洞察。基准已公开:https://openlex3d.github.io/
原文摘要 · Abstract (English)
3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents OpenLex3D, a dedicated benchmark for evaluating 3D open-vocabulary scene representations. OpenLex3D provides entirely new label annotations for scenes from Replica, ScanNet++, and HM3D, which capture real-world linguistic variability by introducing synonymical object categories and additional nuanced descriptions. Our label sets provide 13 times more labels per scene than the original datasets. By introducing an open-set 3D semantic segmentation task and an object retrieval task, we evaluate various existing 3D open-vocabulary methods on OpenLex3D, showcasing failure cases, and avenues for improvement. Our experiments provide insights on feature precision, segmentation, and downstream capabilities. The benchmark is publicly available at: https://openlex3d.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。