系统梳理高效视觉-语言-动作模型的压缩方法,助力机器人在边缘设备实时运行。
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- 按架构、感知特征、动作生成和训练推理四维度归纳效率优化技术
- 针对边缘设备算力与内存限制,提出降低延迟、内存和成本的综合方案
- 适合关注机器人智能部署与轻量化模型的研究者与工程师
视觉-语言-动作(VLA)模型通过将自然语言指令和视觉观测映射为机器人动作,扩展了视觉-语言模型在具身控制中的应用。尽管具备强大能力,现有VLA系统因计算与内存需求巨大,难以在移动机器人等边缘平台实现实时运行。为缓解这一矛盾,本综述系统梳理了提升VLA效率的各类方法,重点降低延迟、内存占用及训练与推理成本。我们从模型架构、感知特征、动作生成和训练/推理策略四个维度分类总结代表性技术,并讨论未来趋势与开放挑战,指明高效具身智能的发展方向。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models extend vision-language models to embodied control by mapping natural-language instructions and visual observations to robot actions. Despite their capabilities, VLA systems face significant challenges due to their massive computational and memory demands, which conflict with the constraints of edge platforms such as on-board mobile manipulators that require real-time performance. Addressing this tension has become a central focus of recent research. In light of the growing efforts toward more efficient and scalable VLA systems, this survey provides a systematic review of approaches for improving VLA efficiency, with an emphasis on reducing latency, memory footprint, and training and inference costs. We categorize existing solutions into four dimensions: model architecture, perception feature, action generation, and training/inference strategies, summarizing representative techniques within each category. Finally, we discuss future trends and open challenges, highlighting directions for advancing efficient embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。