arXiv:2506.09049cs.AIcs.CV2025-06NeurIPS被引 15

用视觉语言模型实现异构智能体协作,支持多视角感知与动态任务规划。

VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning

  • 两阶段框架:先用思维链数据微调视觉语言模型,再通过多层级奖励强化学习
  • 在三个任务层级上均显著优于基线方法,提升协作效率
  • 适用于需要视觉理解的机器人协同场景,如智能仓储、灾害救援

在动态环境中协调多个具身智能体仍是人工智能的核心挑战,需结合感知驱动推理与可扩展的协作策略。尽管近期工作已利用大语言模型进行多智能体规划,少数研究开始探索视觉语言模型(VLM)用于视觉推理,但现有方法对多样化具身形态的支持仍有限。本文提出VIKI-Bench,首个面向具身多智能体协作的分层基准,包含三个结构化层级:智能体激活、任务规划与轨迹感知。该基准涵盖多种机器人形态、多视角视觉观测及结构化监督信号,用于评估基于视觉输入的推理能力。为验证其有效性,我们提出VIKI-R,一种两阶段框架:首先使用思维链标注示范微调预训练视觉语言模型,随后在多层级奖励信号下进行强化学习。大量实验表明,VIKI-R在所有任务层级上均显著优于基线方法。此外,强化学习促使异构智能体间涌现出组合式协作模式。VIKI-Bench与VIKI-R共同提供了一个统一的测试平台与方法,推动具身智能系统中视觉驱动多智能体协作的发展。

原文摘要 · Abstract (English)

Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large language models (LLMs) for multi-agent planning, a few have begun to explore vision-language models (VLMs) for visual reasoning. However, these VLM-based approaches remain limited in their support for diverse embodiment types. In this work, we introduce VIKI-Bench, the first hierarchical benchmark tailored for embodied multi-agent cooperation, featuring three structured levels: agent activation, task planning, and trajectory perception. VIKI-Bench includes diverse robot embodiments, multi-view visual observations, and structured supervision signals to evaluate reasoning grounded in visual inputs. To demonstrate the utility of VIKI-Bench, we propose VIKI-R, a two-stage framework that fine-tunes a pretrained vision-language model (VLM) using Chain-of-Thought annotated demonstrations, followed by reinforcement learning under multi-level reward signals. Our extensive experiments show that VIKI-R significantly outperforms baselines method across all task levels. Furthermore, we show that reinforcement learning enables the emergence of compositional cooperation patterns among heterogeneous agents. Together, VIKI-Bench and VIKI-R offer a unified testbed and method for advancing multi-agent, visual-driven cooperation in embodied AI systems.

多智能体视觉语言模型具身智能强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。