arXiv:2502.15322cs.CVcs.AI2025-02被引 2

用多源元数据增强图像情感分析,提升判断准确性。

SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis

  • 融合图像与文本标签等多源元数据,统一表征输入
  • 自适应加权机制筛选有效元数据,抑制噪声信息
  • 跨模态融合模块实现图文联合推理,适合社交媒体情感分析

随着用户在互联网上越来越多地通过图片表达情绪,图像情感分析受到广泛关注。现有方法主要依赖神经网络提取视觉特征,但对描述图像的元数据(如文字说明、关键词标签)利用不足。本文提出SentiFormer,一种元数据增强的Transformer模型,将多种元数据与图像统一建模。首先获取图像的多源元数据并统一表示;为自适应学习各元数据权重,设计相关性自适应学习模块,突出有效信息、抑制低效内容;进一步构建跨模态融合模块,融合加权后的表示完成情感预测。在三个公开数据集上的实验表明,该方法显著优于现有基准,具备合理性和有效性。

原文摘要 · Abstract (English)

As more and more internet users post images online to express their daily emotions, image sentiment analysis has attracted increasing attention. Recently, researchers generally tend to design different neural networks to extract visual features from images for sentiment analysis. Despite the significant progress, metadata, the data (e.g., text descriptions and keyword tags) for describing the image, has not been sufficiently explored in this task. In this paper, we propose a novel Metadata Enhanced Transformer for sentiment analysis (SentiFormer) to fuse multiple metadata and the corresponding image into a unified framework. Specifically, we first obtain multiple metadata of the image and unify the representations of diverse data. To adaptively learn the appropriate weights for each metadata, we then design an adaptive relevance learning module to highlight more effective information while suppressing weaker ones. Moreover, we further develop a cross-modal fusion module to fuse the adaptively learned representations and make the final prediction. Extensive experiments on three publicly available datasets demonstrate the superiority and rationality of our proposed method.

情感分析图像理解元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。