arXiv:2511.17587cs.LGcs.AI2025-11AAAI被引 3

联合建模情绪与意图,提升表情包回复选择准确率。

Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection

  • 双层次对比框架对齐跨模态情感与意图特征。
  • 在两个公开数据集上优于当前最优模型,显著提升准确率。
  • 适合需要精准理解对话情感与隐含意图的研究者。

表情包广泛用于在线交流以传达情感和隐含意图。表情包回复选择(SRS)任务旨在根据对话上下文选择最合适的表情包。然而,现有方法通常依赖语义匹配,并分别建模情感与意图,当两者不一致时易产生偏差。为此,本文提出情感与意图引导的多模态学习(EIGML),首次实现情感与意图的联合建模,有效减少孤立建模带来的偏差,显著提升选择准确性。具体地,设计双层次对比框架,实现模态内与模态间的一致性对齐;并构建意图-情感引导的多模态融合模块,通过情感引导意图知识选择、意图-情感引导注意力融合及相似度调整匹配机制,逐步整合情感与意图信息,增强对对话的深层理解。在两个公开的SRS数据集上的实验表明,EIGML持续优于现有最优基线,在准确率和情感意图理解方面表现更优。代码已提供于补充材料中。

原文摘要 · Abstract (English)

Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. However, existing methods typically rely on semantic matching and model emotional and intentional cues separately, which can lead to mismatches when emotions and intentions are misaligned. To address this issue, we propose Emotion and Intention Guided Multi-Modal Learning (EIGML). This framework is the first to jointly model emotion and intention, effectively reducing the bias caused by isolated modeling and significantly improving selection accuracy. Specifically, we introduce Dual-Level Contrastive Framework to perform both intra-modality and inter-modality alignment, ensuring consistent representation of emotional and intentional features within and across modalities. In addition, we design an Intention-Emotion Guided Multi-Modal Fusion module that integrates emotional and intentional information progressively through three components: Emotion-Guided Intention Knowledge Selection, Intention-Emotion Guided Attention Fusion, and Similarity-Adjusted Matching Mechanism. This design injects rich, effective information into the model and enables a deeper understanding of the dialogue, ultimately enhancing sticker selection performance. Experimental results on two public SRS datasets show that EIGML consistently outperforms state-of-the-art baselines, achieving higher accuracy and a better understanding of emotional and intentional features. Code is provided in the supplementary materials.

表情包生成多模态学习情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。