用频谱分解分离机器人动作的宏观意图与精细调整,提升操作精度。
Hierarchical Policy Learning via Spectral Decomposition

- 将动作分解为低频全局轨迹和高频精细控制,分步生成
- 在模拟与真实场景中均显著优于基线,尤其在高精度任务上
- 引入人类操作噪声增强数据,提升对错误示范的鲁棒性
本文识别出机器人动作序列中的语义分解结构:任务级运动意图与执行级修正相分离。通过离散余弦变换(DCT)在频域分析动作,发现低频成分捕捉全局运动轨迹,高频成分编码精确的时间、对齐与接触行为。受此启发,我们提出因果频谱策略(CSP),将动作生成建模为因果式的粗粒度到细粒度过程:先由观测与语言预测粗略运动,再基于实际轨迹条件生成精细修正。在仿真与真实世界评估中,CSP 在精度敏感的操控任务上持续优于强基线。此外,我们提出一种类人遥控噪声注入的数据增强方法,使该方法在噪声示范下仍表现出强鲁棒性。
原文摘要 · Abstract (English)
In this paper, we identify a semantic decomposition in robot action sequences, separating task-level motion intent from execution-level refinements. By analyzing actions in the spectral domain using the discrete cosine transform (DCT), we observe that low-frequency components capture global motion trajectories, while high-frequency components encode precise timing, alignment, and contact behaviors. Motivated by this structure, we propose Causal Spectral Policy (CSP), which models action generation as a causal coarse-to-fine process: coarse motion is predicted from observation and language, and fine corrections are generated conditionally on the realized trajectory. Across simulation and real-world evaluations, CSP consistently outperforms strong baselines on precision-sensitive manipulation tasks. Additionally, we propose human-inspired teleoperation noise injection as a data augmentation method, under which our approach demonstrates strong robustness to noisy demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。