arXiv:2512.11239cs.CV2025-12AAAI被引 6

通过跨模态提示提升缺失数据下的多模态情感识别准确率

Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition

  • 用动态提示生成器提取各模态语义线索
  • 在4个数据集上优于7种先进方法,最高提升6.2%准确率
  • 适合处理模态缺失场景的情感分析任务

不完整多模态情感识别(IMER)旨在通过部分观测的多源数据理解人类意图与情绪。尽管多模态数据应提供更丰富信息,但性能差距和模态欠优化问题在数据缺失时尤为突出。为此,我们提出一种新型跨模态提示(ComP)方法,通过增强模态特异性特征并提升各模态表现来实现信息一致性和整体识别精度。具体而言,设计了一个具有动态梯度调制的渐进式提示生成模块,以产生简洁一致的模态语义提示;同时,跨模态知识传播机制利用提示选择性放大模态特征中的一致信息,增强模态输出判别力。此外,引入协调器动态重加权模态输出,作为平衡策略的补充,提升模型有效性。在4个数据集上,针对7种SOTA方法在不同缺失率下的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially observed multi-source data. Although the multi-modal data is expected to provide more abundant information, the performance gap and modality under-optimization problem hinder effective multi-modal learning in practice, and are exacerbated in the confrontation of the missing data. To address this issue, we devise a novel Cross-modal Prompting (ComP) method, which emphasizes coherent information by enhancing modality-specific features and improves the overall recognition accuracy by boosting each modality's performance. Specifically, a progressive prompt generation module with a dynamic gradient modulator is proposed to produce concise and consistent modality semantic cues. Meanwhile, cross-modal knowledge propagation selectively amplifies the consistent information in modality features with the delivered prompts to enhance the discrimination of the modality-specific output. Additionally, a coordinator is designed to dynamically re-weight the modality outputs as a complement to the balance strategy to improve the model's efficacy. Extensive experiments on 4 datasets with 7 SOTA methods under different missing rates validate the effectiveness of our proposed method.

情感识别多模态提示学习缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。