通过锚点融合与语义同步,提升多模态意图识别效果
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
- 用锚点机制融合文本、手势、语调等多模态输入
- 结合大模型生成标签描述,实现跨模态表示同步
- 在多个数据集上表现领先,适合多模态交互研究者
在多模态意图识别(MIR)领域,目标是通过整合语言文本、身体姿态和语调等多种模态来识别人类意图。现有方法难以充分捕捉模态间的内在关联,并忽略意图的对应语义表征。为此,我们提出锚点式多模态嵌入与语义同步框架(A-MESS)。首先设计锚点式多模态嵌入(A-ME)模块,采用锚点嵌入融合机制整合多模态输入;其次构建基于三元组对比学习的语义同步(SS)策略,通过大语言模型生成的标签描述优化多模态表示的对齐。大量实验表明,A-MESS达到当前最优性能,为多模态表征及下游任务提供了重要见解。
原文摘要 · Abstract (English)
In the domain of multimodal intent recognition (MIR), the objective is to recognize human intent by integrating a variety of modalities, such as language text, body gestures, and tones. However, existing approaches face difficulties adequately capturing the intrinsic connections between the modalities and overlooking the corresponding semantic representations of intent. To address these limitations, we present the Anchor-based Multimodal Embedding with Semantic Synchronization (A-MESS) framework. We first design an Anchor-based Multimodal Embedding (A-ME) module that employs an anchor-based embedding fusion mechanism to integrate multimodal inputs. Furthermore, we develop a Semantic Synchronization (SS) strategy with the Triplet Contrastive Learning pipeline, which optimizes the process by synchronizing multimodal representation with label descriptions produced by the large language model. Comprehensive experiments indicate that our A-MESS achieves state-of-the-art and provides substantial insight into multimodal representation and downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。