arXiv:2607.28993cs.ROcs.CV2026-07被引 1

用语义时序建模提升机器人在视觉变化下的操作鲁棒性

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

论文配图:ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
图 1 · 摘自论文原文
  • 引入DINOv3与VAE双空间预测,分离语义与细节动态
  • 在LIBERO上达98.7%成功率,视觉漂移下实测成功率翻倍至61.5%
  • 无需额外预训练或标注,适合真实场景的鲁棒机械臂控制

世界动作模型(WAMs)通过联合建模机器人动作与未来视觉动态成为新范式。但其依赖像素生成的未来监督,易将任务无关视觉内容与动作相关状态混淆,导致视觉分布偏移下性能下降。我们发现‘训练分布幻觉’现象:在视觉变化观测下,模型会幻化出训练域内容而非忠实反映当前场景。对照实验表明,DINOv3特征在视觉变化中更稳定且更保留任务状态区分度,优于Wan-VAE潜变量。为此提出语义-时序世界动作模型(ST-WAM),以DINOv3作为共享语义表示用于未来预测与历史检索,同时保留细粒度VAE动态。其双空间未来专家(DSFE)联合预测未来VAE潜变量与DINO特征,当前锚定意图检索(CAIR)在当前视觉-语言上下文中从近期DINO历史中检索任务相关证据。ST-WAM端到端训练,无需额外具身预训练或任务特定标注,推理时无需显式未来生成。在LIBERO上达到98.7%,在RoboTwin 2.0上达92.8%;相比Fast-WAM,零样本LIBERO-Plus性能提升21.3个百分点,真实世界在视觉偏移下成功率从25.8%提升至61.5%。结果表明,语义-时序建模能有效补充像素生成动态,提升操作鲁棒性。

原文摘要 · Abstract (English)

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.

世界模型机器人操控视觉鲁棒性DINOv3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。