arXiv:2511.06404cs.CV2025-11被引 1

构建首个图文情感标注数据集,揭示信息图如何影响用户情绪。

InfoAffect: Affective Annotations of Infographics in Information Spread

  • 用多模态大模型分析图文内容,融合结果提升情感判断准确率。
  • 创建3500样本数据集,用户评估显示情感一致性达0.608,可靠性高。
  • 适合研究社交媒体传播、情感计算或人机交互的学者使用。

信息图在社交媒体中广泛用于传递复杂信息,但其对用户情绪的影响因缺乏相关数据集而研究不足。为此,我们构建了包含3500个样本的InfoAffect情感标注数据集,结合文本与真实信息图。数据从六个领域收集,并通过预处理、文本优先对齐方法及三项策略确保质量与合规性。我们设计了情感表(Affect Table)规范标注流程。采用五种前沿多模态大语言模型(MLLMs)分析双模态信息,利用互逆排名融合(RRF)算法整合输出,获得稳健的情感判断与置信度。通过两项用户实验验证可用性,并以综合情感一致性指数(CACI)评估数据集,整体得分0.608,表明高准确性。该数据集已公开于https://github.com/bulichuchu/InfoAffect-dataset。

原文摘要 · Abstract (English)

Infographics are widely used in social media to convey complex information, yet how they influence users' affects remains underexplored due to the scarcity of relevant datasets. To address this gap, we introduce a 3.5k-sample affect-annotated InfoAffect dataset, which combines textual content with real-world infographics. We first collected the raw data from six fields and aligned it via preprocessing, the accompanied-text-priority method, and three strategies to guarantee quality and compliance. After that, we constructed an Affect Table to constrain annotation. We used five state-of-the-art multimodal large language models (MLLMs) to analyze both modalities, and their outputs were fused with the Reciprocal Rank Fusion (RRF) algorithm to yield robust affects and confidences. We conducted a user study with two experiments to validate usability and assess the InfoAffect dataset using the Composite Affect Consistency Index (CACI), achieving an overall score of 0.608, which indicates high accuracy. The InfoAffect dataset is available in a public repository at https://github.com/bulichuchu/InfoAffect-dataset.

情感分析信息图多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。