融合语言与深度信息,提升复杂场景下的指代理解准确率
A Multimodal Depth-Aware Method For Embodied Reference Understanding
- 用大模型生成增强数据,结合深度图和决策模块
- 在两个数据集上显著优于现有方法,提升指代识别精度
- 适合需要精准环境交互的机器人应用
具身指代理解需根据语言指令和指向提示,在视觉场景中定位目标物体。尽管已有研究在开放词汇目标检测上取得进展,但在存在多个候选对象的模糊场景中仍表现不佳。为此,我们提出一种新型ERU框架,联合利用基于大语言模型的数据增强、深度图模态以及深度感知决策模块。该设计实现了语言与具身线索的鲁棒融合,提升了复杂或杂乱环境中的消歧能力。在两个数据集上的实验结果表明,该方法显著优于现有基线,实现了更准确可靠的指代检测。
原文摘要 · Abstract (English)
Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。