轻量级多模态模型M3ET提升机器人视觉语言学习效率
M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
- 用Mamba模块和自适应注意力优化跨模态特征融合
- 预训练推理速度提升2.3倍,参数量减少67%
- 适合资源受限的移动机器人部署
近年来,多模态学习在机器人视觉与信息融合中变得至关重要,尤其在复杂环境中理解人类行为。然而,现有方法难以充分利用文本模态,依赖有监督预训练模型,导致在无监督机器人环境中语义提取受限,且存在显著模态损失。同时,这些方法计算开销大,在实际应用中资源消耗高。为此,我们提出多模态增强型Mamba Transformer(M3ET),一种专为移动端设计的轻量级多模态学习模型。通过引入Mamba模块和基于语义的自适应注意力机制,M3ET优化了特征融合、对齐与模态重建。实验表明,M3ET在跨任务性能上有所提升,预训练推理速度提高2.3倍;核心视觉问答(VQA)任务准确率达到0.74,模型参数量减少67%。尽管在环境问答(EQA)任务上表现有限,但其轻量化设计使其非常适合部署于资源受限的机器人平台。
原文摘要 · Abstract (English)
In recent years, multimodal learning has become essential in robotic vision and information fusion, especially for understanding human behavior in complex environments. However, current methods struggle to fully leverage the textual modality, relying on supervised pretrained models, which limits semantic extraction in unsupervised robotic environments, particularly with significant modality loss. These methods also tend to be computationally intensive, leading to high resource consumption in real-world applications. To address these challenges, we propose the Multi Modal Mamba Enhanced Transformer (M3ET), a lightweight model designed for efficient multimodal learning, particularly on mobile platforms. By incorporating the Mamba module and a semantic-based adaptive attention mechanism, M3ET optimizes feature fusion, alignment, and modality reconstruction. Our experiments show that M3ET improves cross-task performance, with a 2.3 times increase in pretraining inference speed. In particular, the core VQA task accuracy of M3ET remains at 0.74, while the model's parameter count is reduced by 0.67. Although performance on the EQA task is limited, M3ET's lightweight design makes it well suited for deployment on resource-constrained robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。