提出频域一致性约束,实现机器人视觉-动作策略的高效一步生成。
FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency
- 通过频域一致性约束,让流模型捕捉动作时序结构。
- 在3个仿真基准上53个任务中优于现有一步生成方法。
- 实测推理频率达93.5 Hz,适合实时机器人系统应用。
基于生成建模的视觉-动作策略在机器人操作中广泛应用,因其能建模多模态动作分布。然而,多步采样的高推理开销限制了其在实时机器人系统中的应用。现有方法借鉴图像生成加速技术,但二者关键差异在于:图像生成通常产生无时间依赖的独立样本,而机器人操作需生成具有连续性与时序一致性的动作轨迹。为此,我们提出FreqPolicy,首次在流模型驱动的视觉-动作策略中引入频域一致性约束。该方法通过强制流路径上不同时刻的动作特征在频域对齐,促进一步动作生成快速收敛至目标分布。同时设计自适应一致性损失,以捕捉机器人操作任务中固有的时序结构变化。我们在3个仿真基准的53个任务上评估,证明其优于现有一步生成器。进一步将该方法集成至视觉-语言-动作(VLA)模型,在LIBERO的40个任务上实现加速且性能无损。此外,实机测试显示推理频率达93.5 Hz,兼具高效与有效。
原文摘要 · Abstract (English)
Generative modeling-based visuomotor policies have been widely adopted in robotic manipulation, attributed to their ability to model multimodal action distributions. However, the high inference cost of multi-step sampling limits its applicability in real-time robotic systems. Existing approaches accelerate sampling in generative modeling-based visuomotor policies by adapting techniques originally developed to speed up image generation. However, a major distinction exists: image generation typically produces independent samples without temporal dependencies, while robotic manipulation requires generating action trajectories with continuity and temporal coherence. To this end, we propose FreqPolicy, a novel approach that first imposes frequency consistency constraints on flow-based visuomotor policies. Our work enables the action model to capture temporal structure effectively while supporting efficient, high-quality one-step action generation. Concretely, we introduce a frequency consistency constraint objective that enforces alignment of frequency-domain action features across different timesteps along the flow, thereby promoting convergence of one-step action generation toward the target distribution. In addition, we design an adaptive consistency loss to capture structural temporal variations inherent in robotic manipulation tasks. We assess FreqPolicy on 53 tasks across 3 simulation benchmarks, proving its superiority over existing one-step action generators. We further integrate FreqPolicy into the vision-language-action (VLA) model and achieve acceleration without performance degradation on 40 tasks of LIBERO. Besides, we show efficiency and effectiveness in real-world robotic scenarios with an inference frequency of 93.5 Hz.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。