首个多语言多模态情感分析基准,评测大模型情感理解能力
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
- 构建跨35语言的多模态情感数据集,覆盖文本、图像、视频
- 包含4类情感任务,提供可微调的数据集与轻量级模型
- 适合评估大模型在跨文化情感理解中的表现
大型语言模型和视觉语言模型(统称LMs)已在自然语言处理和计算机视觉领域取得显著进展,但在情感分析(如情绪识别与情感极性判断)方面的能力仍待深入探索。这一差距主要源于缺乏全面的评估基准以及情感分析任务本身的复杂性。本文提出MMAFFBen,首个面向多语言多模态情感分析的开源基准。该基准涵盖35种语言,支持文本、图像、视频三种模态,覆盖情感极性、情感强度、情绪分类、情绪强度四类核心任务。同时,我们构建了MMAFFIn数据集用于模型微调,并基于其开发了MMAFFLM-3b与MMAFFLM-7b两个轻量级模型。我们对包括GPT-4o-mini在内的多种代表性模型进行了系统评估,全面对比其情感理解能力。项目开源地址:https://github.com/lzw108/MMAFFBen。
原文摘要 · Abstract (English)
Large language models and vision-language models (which we jointly call LMs) have transformed NLP and CV, demonstrating remarkable potential across various fields. However, their capabilities in affective analysis (i.e. sentiment analysis and emotion detection) remain underexplored. This gap is largely due to the absence of comprehensive evaluation benchmarks, and the inherent complexity of affective analysis tasks. In this paper, we introduce MMAFFBen, the first extensive open-source benchmark for multilingual multimodal affective analysis. MMAFFBen encompasses text, image, and video modalities across 35 languages, covering four key affective analysis tasks: sentiment polarity, sentiment intensity, emotion classification, and emotion intensity. Moreover, we construct the MMAFFIn dataset for fine-tuning LMs on affective analysis tasks, and further develop MMAFFLM-3b and MMAFFLM-7b based on it. We evaluate various representative LMs, including GPT-4o-mini, providing a systematic comparison of their affective understanding capabilities. This project is available at https://github.com/lzw108/MMAFFBen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。