用平衡匹配提升视觉语言动作模型的闭环控制性能
$π_0$-EqM: Equilibrium Matching for Closed-Loop Vision-Language-Action Control

- 以平衡匹配替代原有流匹配,实现更灵活的推理深度控制
- 在300步预算下,平均成功率从40.4%提升至50.2%,LIBERO-10达87.0%
- 揭示残差与成功之间的非单调关系,为策略设计提供新视角
当前,视觉语言动作(VLA)模型已成为机器人操作任务中最具潜力的范式,因其具备良好的任务泛化能力。然而,大多数生成式流匹配动作解码器采用固定采样时长,限制了状态相关的计算效率与控制周期间的时序复用。本文提出π₀-EqM,将π₀中的流匹配专家替换为平衡匹配(EqM)解码器,保持上游VLA结构不变。在300步预算下,π₀-EqM使RoboTwin在19个任务上的平均成功率从40.4%提升至50.2%,在LIBERO上保持竞争力,尤其在LIBERO-10任务中达到87.0%。两次阈值扫描发现残差与成功率间存在任务依赖的非单调关系,我们称之为“平稳性-可执行性差距”。结果表明,迭代式VLA控制中的推理深度是策略设计的一部分,并引入能量视角,或可推动跨任务、跨体态的可组合动作生成研究。
原文摘要 · Abstract (English)
Currently, Vision-Language-Action (VLA) models have become the most adopted paradigm for robotic manipulation for its great potential for task generalization. While most generative flow-matching action decoders for VLA control are often deployed with fixed sampling horizons, limiting state-dependent compute and temporal reuse across control cycles. We present $π_0$-EqM, which replaces the flow-matching expert in $π_0$ with an Equilibrium Matching (EqM) decoder while leaving the upstream VLA stack unchanged. Under a matched 300-step budget, $π_0$-EqM improves RoboTwin average success from 40.4% to 50.2% across 19 tasks and remains competitive on LIBERO, with its clearest gain on LIBERO-10 (87.0%). Two threshold scans reveal a task-dependent non-monotonic relation between residual and success, which we term the stationarity--executability gap. The results suggest that inference depth in iterative VLA control is part of policy design and introduce an energy-based VLA perspective that may inform future work on composable action generation across tasks and embodiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。