解决视觉运动策略的长时序连贯性问题,提升动作流畅度。
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy

- 采用频率优化分块与局部锚定流匹配,增强动作连贯性。
- 在多个基准任务上优于现有方法,长程动作更平滑准确。
- 适合需要高精度、长时序控制的机器人操作场景。
视觉运动策略旨在从专家示范中学习复杂操作任务。然而,生成平滑且连贯的轨迹仍具挑战,需平衡近端精度与远端预见性。现有方法多聚焦于块内动作分布优化,常忽略块间一致性,导致块间不连续严重阻碍长时序动作学习。为此,我们提出FocalPolicy,一种具备预见性的视觉运动策略,结合频率优化分块与局部锚定流匹配。引入前瞻性复合目标,在近端动作中监督时间对齐,同时在多未来动作块上正则化频域结构,以提升跨块连贯性。为高效学习复杂动作分布,设计局部锚定采样,提升一致性流匹配训练中的目标信号传播效率。大量实验表明,FocalPolicy优于现有方法,并验证了模块对其他基线的可泛化性。
原文摘要 · Abstract (English)
Visuomotor policies aim to learn complex manipulation tasks from expert demonstrations. However, generating smooth and coherent trajectories remains challenging, as it requires balancing proximal precision with distal foresight. Existing approaches typically focus on optimizing intra-chunk action distributions, often neglecting the inter-chunk coherence. Consequently, inter-chunk discontinuities significantly impede the learning of coherent long-horizon actions. To overcome this limitation and achieve a synergetic balance between precision and foresight, we propose FocalPolicy, a foresight-aware visuomotor policy that combines Frequency-Optimized Chunking with Locally Anchored flow matching. We introduce a foresight composite objective that supervises time-domain alignment within the proximal actions while regularizing frequency-domain structure over multiple future action chunks to improve cross-chunk coherence. To efficiently learn complex action distributions, we design locally anchored sampling to enhance target signal propagation efficiency during consistency flow matching training. Extensive experiments demonstrate that FocalPolicy outperforms existing approaches and confirm the generalizability of our modules to other baselines. Project website: https://focalpolicy.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。