arXiv:2503.07557cs.RO2025-03被引 14

用自动标注提升视觉语言模型的空间推理能力,让机器人更懂社交场景

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning

  • 通过两轮VQA自动标注,结合少量人工监督实现高效空间理解
  • 在感知、预测、推理、动作和解释上比基线模型提升最高16.26%
  • 适合需要精准空间推理的社交机器人导航任务

我们提出一种新方法AutoSpatial,通过结构化空间定位与最小人工监督结合大规模自动标注的视觉问答对,提升视觉语言模型在社交导航任务中的空间理解能力。训练阶段采用分层双轮视觉问答策略,实现对场景的全局与细节双重理解。实验表明,该方法在空间感知、运动预测、思维链推理、最终动作和解释生成五个方面均优于现有最先进模型。评估采用GPT-4o、Gemini 2.0 Flash和Claude 3.5 Sonnet等专家系统提供交叉验证分数,并由人类评估者对模型性能进行相对排序。相比仅使用人工标注数据训练的基线模型,AutoSpatial在感知与预测(最高+10.71%)、推理(最高+16.26%)、动作决策(最高+20.50%)和解释质量(最高+18.73%)上均有显著提升。

原文摘要 · Abstract (English)

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs auto-labeling, our approach tackles the challenge of VLMs' limited spatial understanding in social navigation tasks. By applying a hierarchical two-round VQA strategy during training, AutoSpatial achieves both global and detailed understanding of scenarios, demonstrating more accurate spatial perception, movement prediction, Chain of Thought (CoT) reasoning, final action, and explanation compared to other SOTA approaches. These five components are essential for comprehensive social navigation reasoning. Our approach was evaluated using both expert systems (GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet) that provided cross-validation scores and human evaluators who assigned relative rankings to compare model performances across four key aspects. Augmented by the enhanced spatial reasoning capabilities, AutoSpatial demonstrates substantial improvements by averaged cross-validation score from expert systems in: perception & prediction (up to 10.71%), reasoning (up to 16.26%), action (up to 20.50%), and explanation (up to 18.73%) compared to baseline models trained only on manually annotated data.

社交机器人空间推理视觉问答自动化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。