arXiv:2512.21722cs.RO2025-12被引 5

让机器人在复杂社交场景中同时生成多个合理行动方案

MAction-SocialNav: Multi-Action Socially Compliant Navigation via Reasoning-enhanced Prompt Tuning

  • 用增强推理的提示调优方法处理动作歧义问题
  • 在79个测试样本上实现0.595的决策质量得分,远超GPT-4o
  • 适合需要多策略应变的真实人机交互导航场景

社交合规导航要求机器人在以人类为中心的环境中安全、得体地移动。然而,社交规范常具有模糊性,在同一场景下可能有多种合理动作。现有方法通常假设仅有一个正确动作,限制了对真实世界社交不确定性的处理能力。本文提出MAction-SocialNav,一种高效视觉语言模型,明确应对动作歧义,支持在单一场景中生成多个合理动作。为增强模型推理能力,引入新型元认知提示(MCP)方法。此外,构建了一个多动作社交合规导航数据集,涵盖不同人群密度、室内外环境及双人标注,共789个样本,三轮对话,随机划分为710个训练样本和79个测试样本。设计五项评估指标,衡量高层决策精度、安全性与多样性。大量实验表明,该模型在保持高效率的同时展现优异社交推理性能:决策质量(APG)达0.595,显著优于零样本GPT-4o(0.000)与Claude(0.025);安全对齐度(ER)为0.264,优于GPT-4o(0.642)与Claude(0.668);且实现实时效率(1.524 FPS,超过3倍于基线)。

原文摘要 · Abstract (English)

Socially compliant navigation requires robots to move safely and appropriately in human-centered environments by respecting social norms. However, social norms are often ambiguous, and in a single scenario, multiple actions may be equally acceptable. Most existing methods simplify this problem by assuming a single correct action, which limits their ability to handle real-world social uncertainty. In this work, we propose MAction-SocialNav, an efficient vision language model for socially compliant navigation that explicitly addresses action ambiguity, enabling generating multiple plausible actions within one scenario. To enhance the model's reasoning capability, we introduce a novel meta-cognitive prompt (MCP) method. Furthermore, to evaluate the proposed method, we curate a multi-action socially compliant navigation dataset that accounts for diverse conditions, including crowd density, indoor and outdoor environments, and dual human annotations. The dataset contains 789 samples, each with three-turn conversation, split into 710 training samples and 79 test samples through random selection. We also design five evaluation metrics to assess high-level decision precision, safety, and diversity. Extensive experiments demonstrate that the proposed MAction-SocialNav achieves strong social reasoning performance while maintaining high efficiency, highlighting its potential for real-world human robot navigation. Compared with zero-shot GPT-4o and Claude, our model achieves substantially higher decision quality (APG: 0.595 vs. 0.000/0.025) and safety alignment (ER: 0.264 vs. 0.642/0.668), while maintaining real-time efficiency (1.524 FPS, over 3x faster).

社交导航多动作生成视觉语言模型机器人决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。