不重新训练,让视觉语言动作模型自动纠正动态偏差。
Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

- 提出一种推理时的封闭解修正器,分速度和路径两通道处理动态变化。
- 在动态场景中成功率提升28.8%,混合环境提升25.9%。
- 适合需要快速适配动态任务的机器人控制应用。
视觉-语言-动作(VLA)模型在超越传统控制范式方面展现出卓越的灵活性与泛化能力。然而,多数现有VLA模型基于单帧观测训练,对时间动态结构上存在盲区,导致在非静态场景中性能严重下降,即使在动态数据集上训练或微调也难以避免。现有方法要么需昂贵重训练,要么面临延迟瓶颈与动作块间时间不一致问题。本文提出无训练的速率与路径修正(Pace-and-Path Correction)框架,作为可封装任意分块动作VLA的推理时算子。通过单一二次代价函数的联合最小化,得到正交分解的双通道统一解:速率通道沿规划方向压缩执行,路径通道施加正交空间偏移,共同吸收块内感知到的动力学信息。我们在专为分离运动变量设计的诊断基准MoveBench上评估该方法。实验表明,本框架在动态仅环境与静态-动态混合环境中,相较最先进无训练封装器与动态自适应方法,成功率达28.8%和25.9%的绝对提升。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLAs are trained under a single-frame observation paradigm, which leaves them structurally blind to temporal dynamics. Consequently, these models degrade severely in non-stationary scenarios, even when trained or finetuned on dynamic datasets. Existing approaches either require expensive retraining or suffer from latency bottlenecks and poor temporal consistency across action chunks. We propose Pace-and-Path Correction, a training-free, closed-form inference-time operator that wraps any chunked-action VLA. From a single quadratic cost, joint minimization yields a unified solution that decomposes orthogonally into two distinct channels. The pace channel compresses execution along the planned direction, while the path channel applies an orthogonal spatial offset, jointly absorbing the perceived dynamics within the chunk window. We evaluate our approach on a comprehensive diagnostic benchmark MoveBench designed to isolate motion as the sole controlled variable. Empirical results demonstrate that our framework consistently outperforms state-of-the-art training-free wrappers and dynamic-adaptive methods and improves success rates by up to 28.8% and 25.9% in absolute terms over foundational VLA models in dynamic-only and static-dynamic mixed environments, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。