arXiv:2504.12680cs.AIcs.CV2025-04被引 42

用强化学习让小模型学会空间推理,仅用5000段视频就追平大模型。

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

  • 大视觉模型看图,小语言模型推理,协作完成空间理解。
  • 训练仅需5000个视频样本,推理准确率媲美顶尖多模态模型。
  • 能自发系统分析,适合研究高效推理与模型泛化的新方向。

人类能从连续视觉观测(如第一人称视频)中感知并推理空间关系,但预训练模型如何获得此类高阶推理能力仍不明确。本文提出Embodied-R,一个结合大规模视觉-语言模型(VLMs)用于感知、小规模语言模型(LMs)用于推理的协同框架。通过设计考虑思维与答案逻辑一致性的新型奖励机制,采用强化学习使模型在有限计算资源下实现慢思考能力。仅在5000个具身视频样本上训练后,搭载3B参数量语言模型的Embodied-R在分布内与分布外的具身空间推理任务中均达到与OpenAI-o1、Gemini-2.5-pro等顶尖多模态推理模型相当的性能。该模型还展现出系统性分析和上下文整合等涌现式思维模式。本文进一步探讨了回答长度、在VLM上训练、奖励设计策略以及监督微调(SFT)与强化学习训练后模型泛化差异等研究问题。

原文摘要 · Abstract (English)

Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilities, especially high-level reasoning, remains unclear. This paper introduces Embodied-R, a collaborative framework combining large-scale Vision-Language Models (VLMs) for perception and small-scale Language Models (LMs) for reasoning. Using Reinforcement Learning (RL) with a novel reward system considering think-answer logical consistency, the model achieves slow-thinking capabilities with limited computational resources. After training on only 5k embodied video samples, Embodied-R with a 3B LM matches state-of-the-art multimodal reasoning models (OpenAI-o1, Gemini-2.5-pro) on both in-distribution and out-of-distribution embodied spatial reasoning tasks. Embodied-R also exhibits emergent thinking patterns such as systematic analysis and contextual integration. We further explore research questions including response length, training on VLM, strategies for reward design, and differences in model generalization after SFT (Supervised Fine-Tuning) and RL training.

空间推理强化学习小模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。