从动作标记视角梳理视觉-语言-动作模型,揭示其统一框架与关键设计差异。
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

- 提出动作标记统一框架,将多种VLA模型归为生成逐步具身的动作令牌序列。
- 系统分类8类动作标记形式,涵盖语言描述、代码、轨迹等不同表达方式。
- 适合想理解VLA模型内在机制或规划研究方向的读者参考。
视觉与语言基础模型在多模态理解、推理与生成方面的显著进展,推动了向物理世界扩展智能的努力,催生了视觉-语言-动作(VLA)模型的蓬勃发展。尽管方法各异,我们发现当前VLA模型可统一于一个框架:视觉与语言输入通过一系列VLA模块处理,生成一系列逐步包含更具体、可执行信息的动作令牌,最终输出可执行动作。我们进一步指出,VLA模型的核心设计差异在于动作令牌的构建方式,可分为语言描述、代码、可操作性、轨迹、目标状态、潜在表示、原始动作和推理八类。然而,对动作令牌的理解仍不全面,严重阻碍了有效VLA研发并模糊了未来方向。因此,本综述旨在通过动作标记化视角,系统分类与解析现有VLA研究,提炼各类标记的优劣,并指出改进空间。通过这一系统性分析,我们勾勒出VLA模型演进的整体图景,强调未充分探索但前景广阔的方向,并为未来研究提供指引,期望推动该领域迈向通用智能。
原文摘要 · Abstract (English)
The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence to the physical world, fueling the flourishing of vision-language-action (VLA) models. Despite seemingly diverse approaches, we observe that current VLA models can be unified under a single framework: vision and language inputs are processed by a series of VLA modules, producing a chain of \textit{action tokens} that progressively encode more grounded and actionable information, ultimately generating executable actions. We further determine that the primary design choice distinguishing VLA models lies in how action tokens are formulated, which can be categorized into language description, code, affordance, trajectory, goal state, latent representation, raw action, and reasoning. However, there remains a lack of comprehensive understanding regarding action tokens, significantly impeding effective VLA development and obscuring future directions. Therefore, this survey aims to categorize and interpret existing VLA research through the lens of action tokenization, distill the strengths and limitations of each token type, and identify areas for improvement. Through this systematic review and analysis, we offer a synthesized outlook on the broader evolution of VLA models, highlight underexplored yet promising directions, and contribute guidance for future research, hoping to bring the field closer to general-purpose intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。