arXiv:2504.04991cs.RO2025-04

用小波域多尺度建模+世界先验记忆,让机器人更懂环境、动作更精准。

Wavelet Policy: Imitation Learning in the Scale Domain with World Prior Memory

  • 在小波域分层建模动作,同时注入场景先验记忆
  • 4个仿真+6个真实任务上超越主流基线
  • 轻量设计适合实际机器人部署

传统视觉-运动模仿学习直接在时域预测机器人动作,常缺乏物理场景感知和有效记忆。本文提出Wavelet Policy,一种轻量级模仿学习框架,结合世界先验记忆(WPM)与小波多尺度动作建模。核心思想是将静态背景图像中的持久物理结构编码为紧凑的记忆标记,融合为世界先验标记,在前向传播中注入编码器。基于此条件表示,对时序对齐的潜在动作标记进行小波域分解,并采用单编码器多解码器(SE2MD)架构,分别建模不同时间尺度的潜在分量。最终通过逆小波变换重构子带,并投影为可执行的动作块。为高效学习世界先验,引入世界先验适配损失,促使背景编码器保留持久场景知识,同时保持轻量与稳定。在四个仿真和六个真实机器人操控任务上的实验表明,Wavelet Policy持续优于强基线。结果证明,将尺度域动作建模与世界先验记忆结合,是具身操控的有效且高效方案。

原文摘要 · Abstract (English)

Conventional visuomotor imitation learning usually predicts future robot actions directly in the time domain. Such formulations often have limited physical scene awareness and weak memory. In this work, we propose Wavelet Policy, a lightweight imitation learning framework that combines World Prior Memory (WPM) with wavelet-based multi-scale action modeling. Our key idea is to encode persistent physical scene structure from static background images into compact memory tokens, which are fused into world-prior tokens and injected into the encoder during forward propagation. Based on this memory-conditioned representation, we further perform wavelet-domain decomposition over horizon-aligned latent action tokens and adopt a Single-Encoder Multiple-Decoder (SE2MD) architecture to model latent components at different temporal scales. The resulting latent subbands are reconstructed through inverse wavelet transform and finally projected into executable action chunks. To facilitate efficient world prior learning, we introduce a world-prior adaptation loss, encouraging the background encoder to retain persistent scene knowledge while remaining lightweight and stable. Extensive experiments on four simulated and six real-world robotic manipulation tasks show that Wavelet Policy consistently outperforms strong baselines. These results demonstrate that combining scale-domain action modeling with world-prior memory provides an effective and efficient solution for embodied manipulation.

模仿学习小波分析机器人控制记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。