arXiv:2507.04331cs.RO2025-07ICCV

用小波变换提升长序列任务中的策略学习精度与稳定性。

Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks

  • 引入可学习的小波分解,实现多尺度观察分析
  • 在机器人操作等任务中显著提升策略可靠性
  • 适合需要长期规划的智能体系统研究者

策略学习致力于为具身人工智能系统中的智能体设计基于感知状态的最优行动策略。长时序任务面临复杂动作与观测序列的管理挑战,且存在多种行为模式。小波分析在信号处理中具有优势,可多尺度分解信号以捕捉全局趋势与细节特征。本文提出一种新型小波策略学习框架,利用小波变换增强策略学习能力。该方法采用可学习的多尺度小波分解,实现对观测的精细分析与长期动作规划。通过引入提升方案(lifting scheme),实现高效的多分辨率分析与动作生成。框架在机器人操作、自动驾驶及多机器人协作等多种复杂场景中进行评估,结果表明其显著提升了所学策略的精度与可靠性。

原文摘要 · Abstract (English)

Policy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of actions and observations with multiple modes. Wavelet analysis offers significant advantages in signal processing, notably in decomposing signals at multiple scales to capture both global trends and fine-grained details. In this work, we introduce a novel wavelet policy learning framework that utilizes wavelet transformations to enhance policy learning. Our approach leverages learnable multi-scale wavelet decomposition to facilitate detailed observation analysis and robust action planning over extended sequences. We detail the design and implementation of our wavelet policy, which incorporates lifting schemes for effective multi-resolution analysis and action generation. This framework is evaluated across multiple complex scenarios, including robotic manipulation, self-driving, and multi-robot collaboration, demonstrating the effectiveness of our method in improving the precision and reliability of the learned policy.

策略学习小波变换长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。