arXiv:2509.25037cs.CL2025-09

通过视觉门控融合提升图文情感分析准确率

GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis

  • 设计三路门控LSTM,分别处理语法、语义和跨模态融合
  • 在两个推特数据集上达到最优性能,超越多个基线模型
  • 适合需要精准图文情感理解的电商评论分析场景

基于方面的情感分析(ABSA)已拓展至多模态领域,用户生成内容常包含文本与图像。然而,现有多模态ABSA(MABSA)模型难以过滤噪声视觉信号,且无法有效对齐跨模态中的方面与观点内容。为此,我们提出GateMABSA,一种新型门控多模态架构,整合句法、语义与融合感知的mLSTM。具体而言,GateMABSA引入三类专用mLSTM:Syn-mLSTM用于融入句法结构,Sem-mLSTM强调方面与语义的相关性,Fuse-mLSTM实现选择性多模态融合。在两个基准推特数据集上的大量实验表明,GateMABSA优于多个基线模型。

原文摘要 · Abstract (English)

Aspect-based Sentiment Analysis (ABSA) has recently advanced into the multimodal domain, where user-generated content often combines text and images. However, existing multimodal ABSA (MABSA) models struggle to filter noisy visual signals, and effectively align aspects with opinion-bearing content across modalities. To address these challenges, we propose GateMABSA, a novel gated multimodal architecture that integrates syntactic, semantic, and fusion-aware mLSTM. Specifically, GateMABSA introduces three specialized mLSTMs: Syn-mLSTM to incorporate syntactic structure, Sem-mLSTM to emphasize aspect--semantic relevance, and Fuse-mLSTM to perform selective multimodal fusion. Extensive experiments on two benchmark Twitter datasets demonstrate that GateMABSA outperforms several baselines.

多模态情感分析门控机制图文融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。