用近期动作历史构建先验,让机器人生成策略更稳定高效
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors

- 用近期动作历史构造时间感知先验,替代传统高斯分布
- 在机械臂操作任务中显著提升成功率,路径更平直
- 适用于行为克隆和强化学习,提升采样效率与最终性能
基于扩散和流匹配的生成策略已成为视觉-运动机器人控制的主要范式。本文提出将标准高斯源分布替换为温热先验(WarmPrior),该先验由易获取的近期动作历史构建,可一致提升机械臂操纵任务的成功率。我们发现这一改进源于概率路径显著更平直,类似修正流中的最优传输耦合效果。除标准行为克隆外,WarmPrior还能重塑先验空间中的探索分布,在强化学习中提升样本效率与最终性能。这些结果表明,源分布是生成式机器人控制中一个重要且被忽视的设计维度。
原文摘要 · Abstract (English)
Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the standard Gaussian source distribution with WarmPrior, a simple temporally grounded prior constructed from readily available recent action history, consistently improves success rates on robotic manipulation tasks. We trace this gain to markedly straighter probability paths, echoing the effect of optimal-transport couplings in Rectified Flow. Beyond standard behavior cloning, WarmPrior also reshapes the exploration distribution in prior-space reinforcement learning, improving both sample efficiency and final performance. Collectively, these results identify the source distribution as an important and underexplored design axis in generative robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。