arXiv:2603.19131cs.LGcs.RO2026-03被引 1

重新定义视觉语言动作模型的效率,关注真实机器人执行中的表现。

From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models

  • 用任务完成时间、动作平滑度等系统级指标替代传统计算效率评估。
  • 压缩模型后任务耗时增加,动作质量下降,即使成功率不变。
  • 适合关注机器人实际部署效果的研究者与工程团队参考。

视觉语言动作(VLA)模型通过联合推理视觉、语言和运动模态,使具身智能体能执行更复杂的任务。然而,我们发现当前研究中对‘效率’的定义——如参数量、浮点运算量或解码吞吐率——并不能反映其在机器人平台上的真实性能。在实际执行中,效率由系统级具身行为决定,例如任务完成时间、轨迹平滑度、累计关节转动量和运动能耗。通过对模型压缩、标记稀疏化和动作序列压缩的控制实验,我们观察到:(1) 传统效率指标降低的方法常导致端到端执行成本上升或动作质量下降,尽管任务成功率保持不变;(2) 系统级具身效率指标揭示了学习动作策略中隐藏的性能差异;(3) 常见适配方法如上下文提示或监督微调仅在特定指标上带来轻微提升,如减小抖动或降低动作频率,但可能以延长完成时间等为代价。综上,传统推理效率指标可能忽略具身执行的关键方面。引入具身效率可提供更完整的策略行为与实用性能视图,实现更公平、全面的VLA模型比较。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently enabled embodied agents to perform increasingly complex tasks by jointly reasoning over visual, linguistic, and motor modalities. However, we find that the prevailing notion of ``efficiency'' in current VLA research, characterized by parameters, FLOPs, or token decoding throughput, does not reflect actual performance on robotic platforms. In real-world execution, efficiency is determined by system-level embodied behaviors such as task completion time, trajectory smoothness, cumulative joint rotation, and motion energy. Through controlled studies across model compression, token sparsification, and action sequence compression, we make several observations that challenge common assumptions. (1) Methods that reduce computation under conventional metrics often increase end-to-end execution cost or degrade motion quality, despite maintaining task success rates. (2) System-level embodied efficiency metrics reveal performance differences in the learned action policies that remain hidden under conventional evaluations. (3) Common adaptation methods such as in-context prompting or supervised fine-tuning show only mild and metric-specific improvements in embodied efficiency. While these methods can reduce targeted embodied-efficiency metrics such as jerk or action rate, the resulting gains may come with trade-offs in other metrics, such as longer completion time. Taken together, our results suggest that conventional inference efficiency metrics can overlook important aspects of embodied execution. Incorporating embodied efficiency provides a more complete view of policy behavior and practical performance, enabling fairer and more comprehensive comparisons of VLA models.

具身智能效率评估机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。