测试大模型理解矛盾幽默的能力,发现差距巨大。
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
- 用漫画对比构建新基准,评估模型比较推理能力。
- 顶尖模型在识别关键元素和对比分析上远低于人类。
- 适合研究多模态推理与文化理解的学者参考。
理解幽默——尤其是涉及复杂矛盾叙事、需进行比较推理的幽默——仍是大型视觉语言模型(VLMs)的重大挑战,制约了AI进行类人推理与文化表达的能力。本文通过深入分析以面板对比制造幽默的漫画,探究此问题。我们提出了YesBut(V2)基准,包含1,262张来自多元语言与文化背景的漫画图像,配有详尽标注,涵盖叙事理解的多个层面。利用该基准,我们系统评估了一系列VLMs,涵盖从表面内容理解到深层叙事推理的四项互补任务,尤其关注矛盾元素间的比较推理。大量实验表明,即使最先进的模型也显著落后于人类,常见失败包括视觉感知偏差、关键元素识别错误、比较分析失误及幻觉生成。我们进一步探索了基于文本的训练策略与社会知识增强方法以提升性能。研究不仅揭示了VLMs在文化与创造性表达理解上的核心缺陷,也为发展具备上下文感知、能进行深度叙事理解的模型提供了路径。
原文摘要 · Abstract (English)
Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language models (VLMs). This limitation hinders AI's ability to engage in human-like reasoning and cultural expression. In this paper, we investigate this challenge through an in-depth analysis of comics that juxtapose panels to create humor through contradictions. We introduce the YesBut (V2), a novel benchmark with 1,262 comic images from diverse multilingual and multicultural contexts, featuring comprehensive annotations that capture various aspects of narrative understanding. Using this benchmark, we systematically evaluate a wide range of VLMs through four complementary tasks spanning from surface content comprehension to deep narrative reasoning, with particular emphasis on comparative reasoning between contradictory elements. Our extensive experiments reveal that even the most advanced models significantly underperform compared to humans, with common failures in visual perception, key element identification, comparative analysis and hallucinations. We further investigate text-based training strategies and social knowledge augmentation methods to enhance model performance. Our findings not only highlight critical weaknesses in VLMs' understanding of cultural and creative expressions but also provide pathways toward developing context-aware models capable of deeper narrative understanding though comparative reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。