arXiv:2509.26642cs.RO2025-09被引 28

让机器人通过多感官理解物理世界,提升复杂操作能力。

MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation

  • 用大模型直接处理图像、点云和触觉信号,实现跨模态对齐。
  • 预测未来多感官目标,使模型在真实场景中表现提升12%-24%。
  • 适合需要高精度物理交互的机器人任务,如装配与抓取。

视觉-语言-动作模型(VLAs)通过继承视觉-语言模型并学习动作生成,在机器人操作任务中展现出泛化能力。然而,多数VLA模型仅关注视觉与语言的解析以生成动作,忽略了机器人需在空间-物理世界中感知与交互的需求。为此,本文提出一种多感官语言-动作(MLA)模型,通过协同感知异构感官模态并预测未来多感官目标,促进对物理世界的建模。为增强感知表征,我们设计了一种无需编码器的多模态对齐方案,创新性地将大语言模型本身作为感知模块,通过位置对应关系直接对齐2D图像、3D点云和触觉标记。为进一步提升对物理动态的理解,我们引入未来多感官生成后训练策略,使模型能推理语义、几何及交互信息,从而为动作生成提供更鲁棒的条件。评估表明,MLA模型在复杂、接触丰富的现实任务中分别优于先前最先进的2D与3D VLA方法12%和24%,同时展现出对未见配置的更强泛化能力。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation. Most VLA models focus on interpreting vision and language to generate actions, whereas robots must perceive and interact within the spatial-physical world. This gap highlights the need for a comprehensive understanding of robotic-specific multisensory information, which is crucial for achieving complex and contact-rich control. To this end, we introduce a multisensory language-action (MLA) model that collaboratively perceives heterogeneous sensory modalities and predicts future multisensory objectives to facilitate physical world modeling. Specifically, to enhance perceptual representations, we propose an encoder-free multimodal alignment scheme that innovatively repurposes the large language model itself as a perception module, directly interpreting multimodal cues by aligning 2D images, 3D point clouds, and tactile tokens through positional correspondence. To further enhance MLA's understanding of physical dynamics, we design a future multisensory generation post-training strategy that enables MLA to reason about semantic, geometric, and interaction information, providing more robust conditions for action generation. For evaluation, the MLA model outperforms the previous state-of-the-art 2D and 3D VLA methods by 12% and 24% in complex, contact-rich real-world tasks, respectively, while also demonstrating improved generalization to unseen configurations.

机器人操作多模态感知物理建模语言动作模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。