arXiv:2608.29432cs.RO2026-08

用B样条平滑动作序列,提升长时序视觉语言动作模型的准确率与效率

SMILE: Smooth Motion for Improved Long-Horizon VLA Execution

论文配图:SMILE: Smooth Motion for Improved Long-Horizon VLA Execution
图 1 · 摘自论文原文
  • 通过预测B样条系数生成平滑动作,不改变原有模型结构
  • 在LIBERO上达98.0%成功率,速度提升1.1倍,非边界加速度降78.6%
  • 适用于多种主流VLA模型,实测机械臂动作更稳定、失败更少

视觉语言动作(VLA)模型通过单次调用执行多个动作来降低推理成本,但较长的执行时序常因原始动作块包含抖动和异常值而降低准确性。本文提出SMILE,一种保持架构不变的接口,通过预测B样条系数并解码为平滑动作序列。SMILE仅改变动作表示方式,可在不修改基线模型主干和规模的前提下实现更长的固定时序执行。将SMILE应用于SmolVLA、Evo1、VPP和DAWN,在LIBERO、CALVIN及真实世界实验中均提升了准确率与平均推理效率。SMILE-Evo1在LIBERO上达到98.0%成功率,提速1.1倍;SMILE-VPP在CALVIN上平均动作长度达4.42,提速1.5倍。在匹配执行时序10的情况下,SMILE-SmolVLA使非边界加速度降低78.6%,速度符号变化率下降42.3%。真实xArm测试显示成功率更高,掉落和接触次数更少。结果表明,系数空间的平滑生成是实现高精度、高效率长时序VLA执行的有效路径。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and outliers. We introduce SMILE, an architecture-preserving interface that predicts B-spline coefficients and decodes them into smooth action sequences. SMILE changes only the action representation, enabling longer fixed horizons while retaining each baseline's backbone and model scale. We apply SMILE to SmolVLA, Evo1, VPP, and DAWN, improving accuracy and amortized inference efficiency across LIBERO, CALVIN, and real-world experiments. SMILE-Evo1 reaches 98.0% with a 1.1x speedup on LIBERO, while SMILE-VPP reaches an average length of 4.42 with a 1.5x speedup on CALVIN. At a matched execution horizon of 10, SMILE-SmolVLA reduces non-boundary acceleration by 78.6% and velocity sign-change rate by 42.3%. Real-world xArm tests show higher success, fewer drops, and fewer contacts. These results establish smooth coefficient-space generation as a route to accurate, efficient long-horizon VLA execution. Project page: jongwoopark7978.github.io/smilevla

视觉语言动作动作平滑长时序控制机器人执行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。