arXiv:2507.12695cs.CL2025-07中稿 · ASONAM 2025被引 2

自适应注意力机制提升图文情感分析准确率

AdaptiSent: Context-Aware Adaptive Attention for Multimodal Aspect-Based Sentiment Analysis

  • 通过动态调整图文交互注意力,聚焦上下文相关特征
  • 在推特数据集上各项指标均超越现有模型
  • 适合需要精准理解图文情感的场景

我们提出AdaptiSent,一种用于多模态方面情感分析(MABSA)的新框架,利用自适应跨模态注意力机制,从文本和图像中提升情感分类与方面词提取效果。模型融合动态模态加权与上下文感知注意力,通过关注文本线索与视觉语境的交互,增强情感与方面信息的提取能力。我们在多个基准模型(包括传统文本模型和其他多模态方法)上进行了测试,结果表明,AdaptiSent在标准推特数据集上的精确率、召回率和F1分数均优于现有方法,尤其擅长捕捉对准确分析至关重要的细微跨模态关系。其优势源于根据上下文相关性动态调整关注点的能力,显著提升了各类多模态数据集上的情感分析深度与准确性。AdaptiSent为MABSA设立了新标准,明显超越当前主流方法,尤其在解析复杂多模态信息方面表现突出。

原文摘要 · Abstract (English)

We introduce AdaptiSent, a new framework for Multimodal Aspect-Based Sentiment Analysis (MABSA) that uses adaptive cross-modal attention mechanisms to improve sentiment classification and aspect term extraction from both text and images. Our model integrates dynamic modality weighting and context-adaptive attention, enhancing the extraction of sentiment and aspect-related information by focusing on how textual cues and visual context interact. We tested our approach against several baselines, including traditional text-based models and other multimodal methods. Results from standard Twitter datasets show that AdaptiSent surpasses existing models in precision, recall, and F1 score, and is particularly effective in identifying nuanced inter-modal relationships that are crucial for accurate sentiment and aspect term extraction. This effectiveness comes from the model's ability to adjust its focus dynamically based on the context's relevance, improving the depth and accuracy of sentiment analysis across various multimodal data sets. AdaptiSent sets a new standard for MABSA, significantly outperforming current methods, especially in understanding complex multimodal information.

多模态情感分析自适应注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。