用世界模型提升机器人任务进度评估,更好判断数据质量。
World Value Models for Robotic Manipulation
- 结合世界模型与价值估计,构建能理解时间序列的通用价值模型
- 在标准和非最优轨迹数据上均达到当前最好性能(SOTA)
- 适合用于从混合质量数据中学习机器人策略的场景
通用价值模型在规模化机器人策略学习中至关重要,需具备深度时序理解能力,以结合历史上下文并规划未来结果。然而,现有基于视觉-语言模型(VLM)的机器人价值模型主要依赖静态或时间稀疏的视觉观测预训练,缺乏足够的时序建模能力。相比之下,世界模型天然擅长时序建模与未来规划,是构建可泛化价值函数的理想基础。为此,本文提出世界价值模型(World Value Model, WVM),通过融合世界模型实现准确的任务进度评估,以判断数据质量。在标准基准上,WVM取得当前最优的值序相关性(Value-Order Correlation, VOC)表现。此外,我们引入新基准Suboptimal-Value-Bench,包含800条高保真、人工标注帧的次优轨迹,覆盖多类机器人平台。实验表明,WVM在该基准上仍保持SOTA性能,验证其对专家与次优数据的鲁棒性。将其用于策略学习时,WVM在多种策略提取方法中均提升模拟与真实场景下的操作性能,为混合质量数据学习提供可靠指导。
原文摘要 · Abstract (English)
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal understanding, requiring models to both ground the current belief using historical context and plan over future outcomes. However, most existing robotic value models are built on Vision-Language Model (VLM) backbones that are pretrained primarily on static or temporally sparse visual observations, lacking the requisite temporal modeling capabilities for value estimation. Unlike VLMs, world models naturally excel at temporal modeling and future planning, making them ideal foundations for learning generalizable value functions. Driven by this insight, we marry world models with value estimation to construct a new generalist robotic value model, World Value Model (WVM), that offers accurate task progressions to assess data quality. On standard benchmarks, WVM delivers state-of-the-art (SOTA) Value-Order Correlation (VOC) results. Complementing standard evaluation suites that contains only expert data, we further introduce Suboptimal-Value-Bench, a multi-embodiment benchmark consisting of 800 suboptimal trajectories with high-fidelity, human-labeled frame annotations. Our evaluations show that WVM maintains its SOTA performance on Suboptimal-Value-Bench, establishing its robustness in handling both expert and suboptimal data. When deployed for policy learning, WVM improves manipulation performance across various policy extraction approaches in both simulated and real-world deployment, providing robust guidance for learning from mixed-quality data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。