arXiv:2412.11906cs.CVcs.AI2024-12ACL被引 6

构建多模态笑点理解基准,评估模型真实幽默理解能力

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension

  • 通过改写标题生成同义/反义句,打破文本依赖陷阱
  • 覆盖多领域图像-标题对,设计多样化问题提升评估全面性
  • 提出渐进式提问策略,显著提升模型笑点理解性能

多模态笑点(如图文结合的幽默或讽刺)是在线多媒体平台常见的表达方式。随着多模态大语言模型(MLLMs)快速发展,亟需评估其对这类内容的理解能力。然而现有笑点理解基准存在三大缺陷:1)语言捷径使模型仅依赖文本;2)问题类型单一;3)仅聚焦特定领域(如漫画)。为此,我们提出名为PunchBench的多模态笑点理解基准,旨在实现精准、全面的评估。为提升评估准确性,我们通过修改原始标题生成同义与反义标题,削弱标题中的捷径影响。为实现全面评估,PunchBench整合了来自多个领域的图像-标题对及多样化问题形式。基于此,我们进行了广泛评估,发现当前顶尖的MLLMs与人类在笑点理解上存在显著差距。为提升模型表现,我们提出简单到复杂的问题链(SC-CoQ)策略,让模型先掌握简单问题,再逐步解决复杂问题。SC-CoQ在PunchBench上有效提升了多种MLLM的性能,优于上下文学习和思维链方法。

原文摘要 · Abstract (English)

Multimodal punchlines, which involve humor or sarcasm conveyed in image-caption pairs, are a popular way of communication on online multimedia platforms. With the rapid development of multimodal large language models (MLLMs), it is essential to assess their ability to effectively comprehend these punchlines. However, existing benchmarks on punchline comprehension suffer from three major limitations: 1) language shortcuts that allow models to solely rely on text, 2) lack of question diversity, and 3) narrow focus on a specific domain of multimodal content (e.g., cartoon). To address these limitations, we introduce a multimodal \textbf{Punch}line comprehension \textbf{Bench}mark, named \textbf{PunchBench}, which is tailored for accurate and comprehensive evaluation of punchline comprehension. To enhance the evaluation accuracy, we generate synonymous and antonymous captions by modifying original captions, which mitigates the impact of shortcuts in the captions. To provide a comprehensive evaluation, PunchBench incorporates diverse question formats and image-captions from various domains. On this basis, we conduct extensive evaluations and reveal a significant gap between state-of-the-art MLLMs and humans in punchline comprehension. To improve punchline comprehension, we propose Simple-to-Complex Chain-of-Question (SC-CoQ) strategy, enabling the models to incrementally address complicated questions by first mastering simple ones. SC-CoQ effectively enhances the performance of various MLLMs on PunchBench, surpassing in-context learning and chain-of-thought.

多模态笑点理解基准测试MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。