arXiv:2502.14420cs.ROcs.CV2025-02EMNLP被引 136

提出新框架让AI同时懂视觉、语言和操控,性能显著提升。

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

  • 分阶段对齐训练,先掌握控制再融合多模态信息
  • 在MMMU上表现是之前的6倍,MMStar达47.2%准确率
  • 适合需要统一理解与机器人控制的智能系统研发

人类具备感知、理解并互动物理世界的一体化认知能力。为何大语言模型难以复现这种整体理解?通过对现有视觉-语言-动作模型(VLA)训练范式的系统分析,我们识别出两大挑战:虚假遗忘(机器人训练覆盖关键视觉-文本对齐)和任务干扰(控制与理解任务联合训练时相互削弱)。为此,我们提出ChatVLA,采用分阶段对齐训练,在初步掌握控制后逐步融入多模态数据,并引入专家混合架构以最小化任务干扰。ChatVLA在视觉问答数据集上表现优异,在多模态理解基准上显著超越现有VLA方法。尤其在MMMU上性能达到之前的六倍,于MMStar取得47.2%得分,且参数效率高于ECoT。此外,在25个真实机器人操作任务中,其表现优于OpenVLA等现有VLA方法。结果表明,该统一框架在实现鲁棒多模态理解与高效机器人控制方面具有巨大潜力。

原文摘要 · Abstract (English)

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key challenges: spurious forgetting, where robot training overwrites crucial visual-text alignments, and task interference, where competing control and understanding tasks degrade performance when trained jointly. To overcome these limitations, we propose ChatVLA, a novel framework featuring Phased Alignment Training, which incrementally integrates multimodal data after initial control mastery, and a Mixture-of-Experts architecture to minimize task interference. ChatVLA demonstrates competitive performance on visual question-answering datasets and significantly surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks. Notably, it achieves a six times higher performance on MMMU and scores 47.2% on MMStar with a more parameter-efficient design than ECoT. Furthermore, ChatVLA demonstrates superior performance on 25 real-world robot manipulation tasks compared to existing VLA methods like OpenVLA. Our findings highlight the potential of our unified framework for achieving both robust multimodal understanding and effective robot control.

多模态机器人控制视觉语言模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。