arXiv:2512.22867cs.CVcs.RO2025-12被引 5

构建首个面向城市社交导航的多模态推理数据集,助力智能体理解社会规范。

MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments

  • 采用五步链式思维框架标注感知、预测、推理等环节,明确物理约束与动作空间。
  • 包含10,110个第一人称样本,Qwen3-VL-8B在动作准确率上达77.65%、碰撞率仅6.09%。
  • 适合研究社交导航、具身智能与视觉语言模型的学者使用。

社交合规导航需要对动态行人和物理约束进行结构化推理,以确保安全且可解释的决策。视觉语言模型(VLMs)为此任务提供了良好基础,因其能融合视觉观察与基于语言的社会知识。然而,现有未微调的VLM仍难以可靠理解细微社会规范,因此任务特定微调至关重要。同时,该任务缺乏大规模第一人称视角数据集。为解决上述挑战,我们提出MUSON,一个包含10,110个第一人称样本的多模态数据集,覆盖多样化的室内外社交场景。MUSON采用五步链式思维标注框架,包括感知、预测、推理、行动与解释,并显式建模静态物理约束,采用标准化六动作决策空间。相比现有社交导航数据集,MUSON在推理、动作与解释方面提供一致标注。我们在MUSON上评估了十种代表性小中型VLM。Qwen3-VL-8B表现最优,动作准确率达0.7765,宏观F1得分为0.7490,碰撞率最低为0.0609。结果表明,MUSON是推动社交合规导航发展的有效且可复用基准。数据集已公开于https://github.com/MUSON-dataset/MUSON/releases/tag/v1.0。

原文摘要 · Abstract (English)

Socially compliant navigation requires structured reasoning about dynamic pedestrians and physical constraints to ensure safe and interpretable decisions. Vision-language models (VLMs) provide a promising foundation for this task because they can integrate visual observations with language-based social knowledge. However, existing untuned VLMs still struggle to reliably understand fine-grained social norms, making task-specific fine-tuning essential. At the same time, no large-scale egocentric dataset is available this task. To address these challenges, we introduce MUSON, a multimodal dataset for short-horizon social navigation containing 10,110 egocentric samples collected across diverse indoor and outdoor social scenes. MUSON adopts a structured five-step chain-of-thought annotation framework comprising perception, prediction, reasoning, action, and explanation. It explicitly models static physical constraints and employs a standardized six-action decision space. Compared with existing social-navigation datasets, MUSON provides consistent annotations for reasoning, actions, and explanations. We evaluate ten representative small-to-medium VLMs on MUSON. Qwen3-VL-8B achieves the strongest decision-level performance, attaining the highest action accuracy of 0.7765 and Macro-F1 score of 0.7490, as well as the lowest collision rate of 0.0609. These results demonstrate that MUSON is an effective and reusable benchmark for advancing socially compliant navigation. The dataset is publicly available at https://github.com/MUSON-dataset/MUSON/releases/tag/v1.0.

社交导航多模态视觉语言模型链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。