一拍即合:让机器人听懂指令后一键执行,又快又准。
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation

- 直接建模平均速度场,一步生成动作,告别多步推演。
- 实测推理速度约71Hz,长任务表现超越OpenVLA和π₀。
- 支持弹性时间跨度,让语义指令与物理执行精准对齐。
扩散策略在具身智能中表现优异,但其迭代去噪过程导致延迟高,现有加速方法常牺牲物理一致性。为此,我们提出ElasticFlow——一种无蒸馏、物理一致的一步式策略框架。通过直接建模平均速度场重构平均场理论,实现从噪声到动作的单步映射。针对机器人任务的时间异质性,引入弹性时间跨度机制,显式编码控制粒度,有效克服频谱偏差,实现语义指令与物理执行时长的高效对齐。在LIBERO、CALVIN和RoboTwin等基准上实验表明,ElasticFlow实现约71Hz的高效1-NFE推理,且在长时序任务上优于OpenVLA和π₀等先进方法,展现出高效、鲁棒且语义对齐的控制潜力。
原文摘要 · Abstract (English)
Diffusion policies have demonstrated exceptional performance in embodied AI. However, their iterative denoising process results in high latency, and existing acceleration methods often sacrifice physical consistency. To address this, we propose ElasticFlow, a distillation-free, physics-consistent one-step policy framework. We reconstruct the Mean Field Theory by directly modeling the average velocity field, enabling a direct single-step mapping from noise to action. Addressing the Temporal Heterogeneity of robotic tasks, we introduce the Elastic Time Horizons mechanism. This mechanism effectively overcomes Spectral Bias by explicitly encoding control granularity, achieving efficient alignment between semantic instructions and physical execution horizons. Experiments on benchmarks such as LIBERO, CALVIN, and RoboTwin demonstrate that ElasticFlow achieves efficient 1-NFE inference (approximately 71Hz). Furthermore, it outperforms state-of-the-art methods, including OpenVLA and $π_0$, on long-horizon tasks, highlighting its potential for efficient, robust, and semantically aligned control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。