用双曲空间提升视觉语言动作模型的语义对齐能力
HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models
- 将多模态特征嵌入双曲空间,更好建模视觉与语言的层次结构
- 在多个基准上超越基线,在泛化性和鲁棒性上表现更优
- 适合研究多模态融合、机器人控制与高效模型设计的学者
视觉-语言-动作(VLA)模型在连接多模态感知与机器人控制方面展现出巨大潜力。然而,现有方法通常直接微调预训练视觉-语言模型(VLM),将语义和视觉特征直接输入策略网络,未能充分解决VLA领域独特的语义对齐挑战。本文提出HMVLA,一种新颖的VLA框架,利用视觉与语言中固有的层次结构实现全面的语义对齐。不同于传统在欧几里得空间进行对齐的方法,HMVLA将多模态特征嵌入双曲空间,更有效地建模图像与文本数据中的层次关系。此外,我们引入一种针对语义对齐优化的稀疏门控专家混合(MoE)机制,增强了图像与文本间的多模态理解,同时提升了效率。大量实验表明,HMVLA在准确率和泛化能力上均优于基线方法。我们还通过重构数据集验证其鲁棒性,进一步测试了跨域适应能力。
原文摘要 · Abstract (English)
Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-tuning of pre-trained Vision-Language Models (VLMs), feeding semantic and visual features directly into a policy network without fully addressing the unique semantic alignment challenges in the VLA domain. In this paper, we propose HMVLA, a novel VLA framework that exploits the inherent hierarchical structures in vision and language for comprehensive semantic alignment. Unlike traditional methods that perform alignment in Euclidean space, our HMVLA embeds multimodal features in hyperbolic space, enabling more effective modeling of the hierarchical relationships present in image text data. Furthermore, we introduce a sparsely gated Mixture of Experts (MoE) mechanism tailored for semantic alignment, which enhances multimodal comprehension between images and text while improving efficiency. Extensive experiments demonstrate that HMVLA surpasses baseline methods in both accuracy and generalization. In addition, we validate its robustness by reconstructing datasets to further test cross domain adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。