arXiv:2506.10011cs.MMcs.AI2025-06IJCAI被引 8

用小波分析非语言信息,提升多模态意图识别准确率

WDMIR: Wavelet-Driven Multimodal Intent Recognition

  • 通过小波变换在频域同步分解视频音频特征
  • 在MIntRec数据集上准确率提升1.13%,细微情绪识别增0.41%
  • 适合做情感分析与多模态交互系统的研究者

多模态意图识别(MIR)旨在融合视频、音频和文本模态的言语与非言语信息,精准理解用户意图。现有方法多侧重文本分析,常忽略非言语线索中的丰富语义。本文提出一种新型小波驱动的多模态意图识别框架(WDMIR),通过非言语信息的频域分析增强意图理解。具体包括:(1) 小波驱动融合模块,在频域同步分解与整合视频-音频特征,实现对时序动态的细粒度分析;(2) 跨模态交互机制,促进从双模态到三模态特征的渐进式增强,有效弥合言语与非言语信息间的语义鸿沟。在MIntRec数据集上的大量实验表明,该方法达到当前最优性能,准确率领先先前方法1.13%。消融实验进一步验证,小波驱动融合模块显著提升非言语源的语义信息提取能力,在分析细微情绪线索时准确率提升0.41%。

原文摘要 · Abstract (English)

Multimodal intent recognition (MIR) seeks to accurately interpret user intentions by integrating verbal and non-verbal information across video, audio and text modalities. While existing approaches prioritize text analysis, they often overlook the rich semantic content embedded in non-verbal cues. This paper presents a novel Wavelet-Driven Multimodal Intent Recognition(WDMIR) framework that enhances intent understanding through frequency-domain analysis of non-verbal information. To be more specific, we propose: (1) a wavelet-driven fusion module that performs synchronized decomposition and integration of video-audio features in the frequency domain, enabling fine-grained analysis of temporal dynamics; (2) a cross-modal interaction mechanism that facilitates progressive feature enhancement from bimodal to trimodal integration, effectively bridging the semantic gap between verbal and non-verbal information. Extensive experiments on MIntRec demonstrate that our approach achieves state-of-the-art performance, surpassing previous methods by 1.13% on accuracy. Ablation studies further verify that the wavelet-driven fusion module significantly improves the extraction of semantic information from non-verbal sources, with a 0.41% increase in recognition accuracy when analyzing subtle emotional cues.

多模态意图识别小波变换情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。