首份高效视觉-语言-动作模型综述,系统梳理三类优化路径。
A Survey on Efficient Vision-Language-Action Models
- 构建统一分类框架,覆盖模型、训练、数据三方面效率优化
- 归纳三大核心方向:高效架构、低耗训练、智能数据采集
- 适合关注机器人智能与多模态模型轻量化的研究者参考
视觉-语言-动作模型(VLAs)是具身智能的重要前沿,旨在连接数字知识与物理世界交互。尽管性能出色,但基础VLAs受限于大规模架构带来的巨大计算与数据需求。近年来虽涌现出大量提升效率的研究,但领域缺乏统一整合框架。本文首次全面综述高效视觉-语言-动作模型(Efficient VLAs),覆盖从模型到训练再到数据的全链条流程。提出统一分类体系,将现有技术归为三大支柱:(1) 高效模型设计,聚焦高效架构与模型压缩;(2) 高效训练,降低学习过程中的计算负担;(3) 高效数据收集,解决机器人数据获取与利用的瓶颈。通过深入分析前沿方法,本文不仅为社区提供基础参考,还总结代表性应用,明确关键挑战,并规划未来研究路线。项目页面持续更新:https://evla-survey.github.io/
原文摘要 · Abstract (English)
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. While a surge of recent research has focused on enhancing VLA efficiency, the field lacks a unified framework to consolidate these disparate advancements. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey not only establishes a foundational reference for the community but also summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。