通过动态对齐提升机器人策略泛化能力,尤其在视觉干扰下表现更优。
Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- 策略与动态模型相互校正,实现动作生成时的自纠错
- 在真实机器人任务中优于基线方法,尤其在分布外场景表现稳定
- 无需大规模预训练,适用于复杂操控任务
机器人行为克隆方法因数据支持有限而泛化能力差。现有基于视频预测的方法虽能从大规模数据中学习时空表征,但其动态模型不区分控制输入,难以用于精确操作任务,且依赖大体量预训练数据。本文提出动态对齐流匹配策略(DAP),将动态预测融入策略学习。该方法设计新架构,使策略与动态模型在动作生成过程中互为反馈,实现自校正与泛化提升。实证结果表明,该方法在真实机器人操纵任务中性能超越基线,在分布外场景(如视觉干扰、光照变化)下仍保持强鲁棒性。
原文摘要 · Abstract (English)
Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction models have shown promising results by learning rich spatiotemporal representations from large-scale datasets. However, these models learn action-agnostic dynamics that cannot distinguish between different control inputs, limiting their utility for precise manipulation tasks and requiring large pretraining datasets. We propose a Dynamics-Aligned Flow Matching Policy (DAP) that integrates dynamics prediction into policy learning. Our method introduces a novel architecture where policy and dynamics models provide mutual corrective feedback during action generation, enabling self-correction and improved generalization. Empirical validation demonstrates generalization performance superior to baseline methods on real-world robotic manipulation tasks, showing particular robustness in OOD scenarios including visual distractions and lighting variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。