构建多物体情感分析基准,评估大模型理解复杂图像能力
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
- 设计多物体情感分析数据集,支持独立判断每个物体情绪
- 实验发现模型对远距离物体情感识别准确率显著下降
- 适合研究视觉-语言模型在复杂场景下语义理解的学者
多模态大语言模型(MLLMs)在视觉问答、图像描述和情绪识别等高层语义任务中取得显著进展。然而,当前缺乏针对多物体情感分析的标准评估基准,该任务是语义理解的关键。为此,我们提出MOSABench,一个专为多物体情感分析设计的新基准数据集,包含约1,000张含多个对象的图像,要求模型独立评估每个对象的情感,以反映真实世界复杂性。MOSABench的核心创新包括基于距离的目标标注、用于标准化输出的后处理流程以及改进的评分机制。实验显示,部分模型如mPLUG-owl和Qwen-VL2能有效关注情感相关特征,但其他模型存在注意力分散问题,且随着物体间空间距离增加,性能明显下降。本研究揭示了提升MLLM在复杂多对象情感分析任务中准确性的必要性,并确立MOSABench作为推动该领域发展的基础工具。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. However, despite advancements, there remains a lack of standardized benchmarks for evaluating MLLMs performance in multi-object sentiment analysis, a key task in semantic understanding. To address this gap, we introduce MOSABench, a novel evaluation dataset designed specifically for multi-object sentiment analysis. MOSABench includes approximately 1,000 images with multiple objects, requiring MLLMs to independently assess the sentiment of each object, thereby reflecting real-world complexities. Key innovations in MOSABench include distance-based target annotation, post-processing for evaluation to standardize outputs, and an improved scoring mechanism. Our experiments reveal notable limitations in current MLLMs: while some models, like mPLUG-owl and Qwen-VL2, demonstrate effective attention to sentiment-relevant features, others exhibit scattered focus and performance declines, especially as the spatial distance between objects increases. This research underscores the need for MLLMs to enhance accuracy in complex, multi-object sentiment analysis tasks and establishes MOSABench as a foundational tool for advancing sentiment analysis capabilities in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。