arXiv:2607.06988cs.ROcs.AI2026-07被引 3

看人操作视频就能让机器人学会新任务,无需额外训练。

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

论文配图:WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
图 1 · 摘自论文原文
  • 用自监督预测把人类视频融入冻结模型的记忆中
  • 测试时仅需人类视频即可适配,性能超越现有方法
  • 适合快速部署新任务,尤其无标注数据场景

让机器人基础模型(RFM)适应新任务或用户偏好行为仍具挑战,通常需额外机器人示范、任务特定微调或长上下文条件。我们提出WAM-TTT,一种基于原始人类视频在测试时调整世界-动作模型的框架。不同于将人类视频视为待模仿的轨迹,WAM-TTT通过自监督视频预测,将人类视频信息注入冻结的WAM模型中的轻量级自适应记忆。为使该记忆可用于控制,引入元训练阶段,利用成对的人机数据与键值记忆重建目标,对齐人类示范与机器人行为。测试时仅需未标注的人类视频即可更新记忆,预训练的WAM保持冻结。该方法实现高效可复用的引导,无需机器人动作、人类侧标注或任务特定微调,同时保留基础模型泛化能力。大量实验表明,WAM-TTT在多样化操控任务和泛化设置下,始终优于上下文人类视频条件基线。

原文摘要 · Abstract (English)

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.

机器人视频引导测试时训练自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。