让机器人视觉模型从看热闹变点评高手,自动判断操作进度
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
- 用强化学习驱动模型生成思考链,主动评估任务进展
- 70亿参数模型误差比基线降低50%,超越720亿参数通用模型
- 零样本检测复杂失败场景,对齐顶尖闭源模型表现
长时序机器人操作中的准确过程监督仍是关键挑战。当前视频多模态大模型主要采用监督微调范式,仅能被动识别正在进行的事件,无法评估当前状态与最终目标的差距。本文提出PRIMO R1(Process Reasoning Induced Monitoring)框架,一个70亿参数的模型,将视频多模态大模型转变为积极的“批评者”。我们利用基于结果的强化学习,激励模型生成显式的思维链以进行进展估计。此外,我们的架构通过显式锚定初始状态与当前状态图像,构建结构化的时序输入。在自建的PRIMO数据集与基准测试支持下,广泛实验涵盖多种域内环境及域外真实人形机器人场景,结果表明PRIMO R1达到业界领先性能。定量结果显示,该70亿参数模型相较专业推理基线,平均绝对误差降低50%;其性能显著优于720亿参数量级的一般多模态大模型。同时,PRIMO R1在困难故障检测任务中展现出强零样本泛化能力,在RoboFail基准上达到67.0%准确率,超越OpenAI o1等闭源模型6.0个百分点。
原文摘要 · Abstract (English)
Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video MLLMs, trained primarily under a Supervised Fine-Tuning (SFT) paradigm, function as passive "Observers" that recognize ongoing events rather than evaluating the current state relative to the final task goal. In this paper, we introduce PRIMO R1 (Process Reasoning Induced Monitoring), a 7B framework that transforms video MLLMs into active "Critics". We leverage outcome-based Reinforcement Learning to incentivize explicit Chain-of-Thought generation for progress estimation. Furthermore, our architecture constructs a structured temporal input by explicitly anchoring the video sequence between initial and current state images. Supported by the proposed PRIMO Dataset and Benchmark, extensive experiments across diverse in-domain environments and out-of-domain real-world humanoid scenarios demonstrate that PRIMO R1 achieves state-of-the-art performance. Quantitatively, our 7B model achieves a 50% reduction in the mean absolute error of specialized reasoning baselines, demonstrating significant relative accuracy improvements over 72B-scale general MLLMs. Furthermore, PRIMO R1 exhibits strong zero-shot generalization on difficult failure detection tasks. We establish state-of-the-art performance on RoboFail benchmark with 67.0% accuracy, surpassing closed-source models like OpenAI o1 by 6.0%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。