让多个智能体从不同视角协同推理世界,提升机器人协作感知能力。
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
- 分两阶段训练,结合思维链与强化学习实现跨视角推理
- 在三个任务上超越主流模型,全局计数准确率提升12.3%
- 适合研究多机器人协作、空间推理与视觉语言融合的学者
理解分布式、局部视角下的世界是具身多智能体系统的核心挑战。每个智能体仅能通过自我中心视角观察环境,常受遮挡和模糊影响。为此,我们提出Ego-to-World(E2W)基准,评估视觉语言模型在三类任务中的跨视角融合能力:(i) 全局计数,(ii) 关系位置推理,(iii) 需要预测视图特定图像坐标的动作导向抓取。为应对该场景,我们提出CoRL框架,采用思维链监督微调与基于组相对策略优化的强化学习相结合。其核心组件——跨视图空间奖励(CVSR)通过将推理步骤与视觉证据关联,提供密集的任务对齐反馈,确保跨视图实体一致识别,并引导模型得出正确最终预测。E2W实验表明,CoRL在推理与感知对齐指标上持续优于强基线模型;消融实验进一步验证了CVSR各组件的必要性。此外,CoRL可泛化至外部空间推理基准,并在配备校准多相机阵列的真实世界多机器人操作中实现有效跨视图定位与成功抓放执行。E2W与CoRL共同构建了从分布式自我中心观测中学习世界中心场景理解的坚实基础,推动协同具身AI发展。
原文摘要 · Abstract (English)
Understanding the world from distributed, partial viewpoints is a fundamental challenge for embodied multi-agent systems. Each agent perceives the environment through an ego-centric view that is often limited by occlusion and ambiguity. To study this problem, we introduce the Ego-to-World (E2W) benchmark, which evaluates a vision-language model's ability to fuse heterogeneous viewpoints across three tasks: (i) global counting, (ii) relational location reasoning, and (iii) action-oriented grasping that requires predicting view-specific image coordinates. To address this setting, we propose CoRL, a two-stage framework that combines Chain-of-Thought supervised fine-tuning with reinforcement learning using Group-Relative Policy Optimization. Its core component, the Cross-View Spatial Reward (CVSR), provides dense task-aligned feedback by linking reasoning steps to visual evidence, ensuring coherent cross-view entity resolution, and guiding the model toward correct final predictions. Experiments on E2W show that CoRL consistently surpasses strong proprietary and open-source baselines on both reasoning and perception-grounding metrics, while ablations further confirm the necessity of each CVSR component. Beyond that, CoRL generalizes to external spatial reasoning benchmarks and enables effective real-world multi-robot manipulation with calibrated multi-camera rigs, demonstrating cross-view localization and successful grasp-and-place execution. Together, E2W and CoRL provide a principled foundation for learning world-centric scene understanding from distributed, ego-centric observations, advancing collaborative embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。