arXiv:2606.09134cs.ROcs.AI2026-06中稿 · ICRA

用大模型零样本自动把3D场景物体映射到知识图谱,准确率超传统方法。

From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs

论文配图:From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs
图 1 · 摘自论文原文
  • 用大模型零样本匹配场景物体与知识图谱类别,无需训练
  • 描述性名称下准确率达90%-96%,缩写名下49%-89%
  • 适合机器人场景理解与自动化知识构建任务

从3D仿真场景构建知识图谱对机器人任务推理至关重要,但核心瓶颈——将场景物体映射到正式本体类——仍依赖人工维护的词典,缺乏泛化能力。本文研究大语言模型(LLMs)能否在Universal Scene Description(USD)场景中实现零样本、免训练的本体标注。在包含125个物体的厨房场景与SOMA-HOME本体上,当使用描述性名称时,LLMs达到90%-96%的精确匹配准确率;使用缩写名称时为49%-89%,显著优于词典与嵌入基线。在完全匿名名称下,通过上下文增强提示可恢复至最高48%。特征消融实验表明,LLMs主要依赖场景图中的语义线索(同级名称与父路径);若隐藏这些线索,准确率降至0-6%,仅几何信息贡献4-17%。

原文摘要 · Abstract (English)

Constructing knowledge graphs from 3D simulation scenes is essential for robot task reasoning, but the key bottleneck, grounding scene objects to formal ontology classes, still relies on manually curated dictionaries that are brittle and do not generalize across assets. We investigate whether large language models (LLMs) can automate this grounding step for Universal Scene Description (USD) scenes as a zero-shot, training-free alternative. On a kitchen scene (125 objects) with SOMA-HOME Ontology, LLMs achieve 90-96% exact-match accuracy with descriptive names and 49-89% with abbreviated names, substantially outperforming dictionary and embedding baselines. Under fully opaque names, context-augmented prompting recovers up to 48%. Feature ablation reveals that LLMs primarily exploit semantic cues in the scene graph (sibling names and parent paths); anonymizing these cues reduces accuracy to 0-6%, while geometry alone yields only 4-17%.

知识图谱大模型3D场景零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。