评测大模型看网络漫画的幽默理解能力,发现差距明显。
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
- 构建2800张多格漫画数据集,标注幽默和叙事顺序。
- 顶尖模型在分镜排序上仅61%准确率,远低于人类表现。
- 适合研究多模态推理、社交智能与幽默理解的学者。
理解幽默是社交智能的核心,但对大型多模态模型(LMMs)仍是重大挑战。我们提出PixelHumor,一个包含2,800张标注多格漫画的基准数据集,用于评估LMMs对多模态幽默及叙事序列的理解能力。对前沿LMMs的实验显示显著差距:例如,顶级模型在分镜排序任务中仅达到61%准确率,远低于人类表现。这凸显当前模型在整合视觉与文本线索以实现连贯叙事和幽默理解方面的关键局限。PixelHumor通过提供严谨的评估框架,旨在推动具备自然、社会感知能力的LMMs发展。
原文摘要 · Abstract (English)
Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。