arXiv:2509.15540cs.CVcs.CL2025-09中稿 · WWW 2026

通过视觉非语言线索提升欲望、情绪和情感识别效果

Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues

  • 构建双向对称多模态框架,实现图文细粒度对齐
  • 融合高低分辨率图像特征,提升意图相关视觉表征能力
  • 在多个任务上实现性能突破,适合多模态情感分析研究者

多模态欲望理解是情感计算中一个新兴但研究不足的任务,旨在从视觉与文本线索中推断人类意图,广泛应用于社交媒体分析。现有方法主要依赖语言信息,忽视了图像中的非语言线索。为此,我们提出对称双向多模态学习框架SyDES,用于欲望、情绪与情感识别。核心在于实现文本与图像模态间的双向细粒度对齐。具体而言,采用混合尺度图像策略,结合低分辨率图像的全局上下文与高分辨率子图上的掩码图像建模(MIM)提取局部细节,有效捕捉意图相关的视觉表示。进一步设计对称交叉模态解码器,包括文本引导的图像解码器和图像引导的文本解码器,实现模态间相互重建与优化,促进深度交互。同时,设计专用损失函数以协调MIM目标与模态对齐目标之间的潜在冲突。在MSED基准上的大量实验表明,该方法取得新最佳性能,欲望理解的F1分数提升1.1%。情绪与情感识别也持续获得增益,验证了其泛化能力及利用非语言线索的必要性。代码已开源:https://github.com/especiallyW/SyDES。

原文摘要 · Abstract (English)

Multimodal desire understanding, a task closely related to both emotion and sentiment that aims to infer human intentions from visual and textual cues, is an emerging yet underexplored task in affective computing with applications in social media analysis. Existing methods for related tasks predominantly focus on mining verbal cues, often overlooking the effective utilization of non-verbal cues embedded in images. To bridge this gap, we propose a Symmetrical Bidirectional Multimodal Learning Framework for Desire, Emotion, and Sentiment Recognition (SyDES). The core of SyDES is to achieve bidirectional fine-grained modal alignment between text and image modalities. Specifically, we introduce a mixed-scaled image strategy that combines global context from low-resolution images with fine-grained local features via masked image modeling (MIM) on high-resolution sub-images, effectively capturing intention-related visual representations. Then, we devise symmetrical cross-modal decoders, including a text-guided image decoder and an image-guided text decoder, which enable mutual reconstruction and refinement between modalities, facilitating deep cross-modal interaction. Furthermore, a set of dedicated loss functions is designed to harmonize potential conflicts between the MIM and modal alignment objectives during optimization. Extensive evaluations on the MSED benchmark demonstrate the superiority of our approach, which establishes a new state-of-the-art performance with 1.1% F1-score improvement in desire understanding. Consistent gains in emotion and sentiment recognition further validate its generalization ability and the necessity of utilizing non-verbal cues. Our code is available at: https://github.com/especiallyW/SyDES.

多模态情感识别视觉理解非语言线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。