arXiv:2509.22729cs.CLcs.AI2025-09中稿 · presentation at th…被引 1

动态融合语音与文本,提升情感分析准确率

Multi-Modal Sentiment Analysis with Dynamic Attention Fusion

  • 用自适应注意力动态加权语音和文本模态
  • 在大型基准上实现更高F1分数,降低预测误差
  • 无需微调预训练模型,适合情感计算应用

传统情感分析长期局限于纯文本任务,忽视了语调、语速等非语言线索对真实情感意图的关键作用。我们提出轻量级的动态注意力融合(DAF)框架,将冻结的预训练语言模型文本嵌入与语音编码器提取的声学特征相结合,通过自适应注意力机制对每条语句的模态进行动态加权。在不微调底层编码器的前提下,DAF模型在大规模多模态基准上持续优于静态融合和单模态基线。实验显示显著提升的F1分数与更低的预测误差,并通过多种消融研究验证了动态加权策略对复杂情感输入建模的重要性。该方法有效整合语言与非语言信息,为情感预测提供更鲁棒的基础,对情绪识别、心理健康评估及人机自然交互等情感计算应用具有广泛影响。

原文摘要 · Abstract (English)

Traditional sentiment analysis has long been a unimodal task, relying solely on text. This approach overlooks non-verbal cues such as vocal tone and prosody that are essential for capturing true emotional intent. We introduce Dynamic Attention Fusion (DAF), a lightweight framework that combines frozen text embeddings from a pretrained language model with acoustic features from a speech encoder, using an adaptive attention mechanism to weight each modality per utterance. Without any finetuning of the underlying encoders, our proposed DAF model consistently outperforms both static fusion and unimodal baselines on a large multimodal benchmark. We report notable gains in F1-score and reductions in prediction error and perform a variety of ablation studies that support our hypothesis that the dynamic weighting strategy is crucial for modeling emotionally complex inputs. By effectively integrating verbal and non-verbal information, our approach offers a more robust foundation for sentiment prediction and carries broader impact for affective computing applications -- from emotion recognition and mental health assessment to more natural human computer interaction.

情感分析多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。