arXiv:2511.05936cs.ROcs.AI2025-11AAAI被引 3

梳理10大关键挑战,推动视觉语言动作模型落地应用

10 Open Challenges Steering the Future of Vision-Language-Action Models

论文配图:10 Open Challenges Steering the Future of Vision-Language-Action Models
图 1 · 摘自论文原文
  • 提出10个核心发展瓶颈,覆盖多模态与推理等方向
  • 强调跨机器人泛化、安全性和人机协同等实际需求
  • 适合关注具身智能与AI落地的科研人员参考

由于能够遵循自然语言指令,视觉语言动作(VLA)模型在具身人工智能领域日益普及,继大型语言模型(LLMs)和视觉语言模型(VLMs)成功之后。本文探讨了VLA模型发展中面临的10个主要里程碑:多模态、推理、数据、评估、跨机器人动作泛化、效率、全身协调、安全性、智能体以及与人类的协作。此外,我们还讨论了空间理解、世界动态建模、后训练和数据合成等新兴趋势,旨在实现这些目标。通过这些讨论,我们希望引起对可能加速VLA模型广泛应用的研究方向的关注。

原文摘要 · Abstract (English)

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread success of their precursors -- LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing development of VLA models -- multimodality, reasoning, data, evaluation, cross-robot action generalization, efficiency, whole-body coordination, safety, agents, and coordination with humans. Furthermore, we discuss the emerging trends of using spatial understanding, modeling world dynamics, post training, and data synthesis -- all aiming to reach these milestones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability.

具身智能多模态智能体人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。