一个能自我进化、自主学习抓取的通用世界模型。
Motus2: A Self-Evolving General World Model for Dexterous Manipulation

- 用统一模型实现感知、预测、决策与评估闭环,支持策略迭代。
- 通过多视角数据和机器人实操数据提升模型泛化能力。
- 适合研究具身智能、机器人操控与自进化系统的研究者。
通用具身智能体应在一个统一系统中完成感知、预测、行动、评估与改进。世界模型在构建此类智能体方面展现出巨大潜力,但现有模型通常仅在世界模拟器后附加动作输出头,未能形成闭环决策与学习回路以优化策略。本文提出 Motus2,一种面向灵巧操作的自进化通用世界模型。该模型通过模型扩展与数据扩展双重路径推进:在模型层面,采用共享权重的单一模型,集成策略(世界-动作模型)、模拟器(条件动作世界模型)与评估器(价值模型)三种控制接口;策略生成动作片段,模拟器预测视觉结果,评估器判断预测成效,三者构成策略优化的闭环。该机制利用专家示范学习动作,失败与次优交互则用于动力学建模与价值学习。在数据层面,从大规模单目内视角数据扩展至同步双目内视角数据,并引入机器人轨迹与人机对齐数据进行领域适配。此外,Motus2还探索了全局自回归与混合记忆的滑动窗口结构,加入触觉反馈实现接触感知控制,并部署于具备双目视觉、双臂、双灵巧手及触觉传感的类生物仿生平台上。整体而言,内视角数据扩展与闭环世界模型扩展共同提供了一条通往自进化灵巧操作的通用路径。
原文摘要 · Abstract (English)
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。