arXiv:2507.21476cs.CLcs.AI2025-07被引 8

用幽默漫画评测大模型的非理工推理能力,发现逻辑能力可跨领域迁移。

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

  • 构建幽默漫画基准,要求模型解释笑点背后的逻辑关联。
  • 顶尖模型在幽默理解上表现良好,且理工训练数据也能有效迁移。
  • 增加思考令牌数对部分模型有帮助,但效果因模型而异。

我们提出HumorBench,一个用于评估大语言模型(LLMs)理解与解释漫画标题中复杂幽默能力的基准。随着推理模型在数学与科学领域趋于饱和,对非理工领域智能的新挑战性评估至关重要。文本类幽默的理解本质上依赖推理,需识别漫画/标题中概念与外部文化引用、双关语等之间的关联。HumorBench包含约300个来自《纽约客》漫画比赛和Cartoonstock.com的独特漫画-标题配对,配有专家标注的评估标准,明确关键笑点要素。通过模型对幽默的解释及其对笑点要素的识别能力进行评估。为在该任务上表现优异,模型需形成并检验概念间关联假设,必要时回溯初始理解以得出最合理解释。对当前主流模型的广泛测试揭示三个关键发现:(1) 模型在理工推理上的进展可有效迁移至幽默理解;(2) 仅在理工数据上训练的模型在HumorBench上仍表现良好,表明推理能力具有强迁移性;(3) 通过增加思考令牌预算进行测试时扩展,在不同模型上效果不一,呈现混合结果。

原文摘要 · Abstract (English)

We present HumorBench, a benchmark designed to evaluate large language models' (LLMs) ability to reason about and explain sophisticated humor in cartoon captions. As reasoning models increasingly saturate existing benchmarks in mathematics and science, novel and challenging evaluations of model intelligence beyond STEM domains are essential. Reasoning is fundamentally involved in text-based humor comprehension, requiring the identification of connections between concepts in cartoons/captions and external cultural references, wordplays, and other mechanisms. HumorBench includes approximately 300 unique cartoon-caption pairs from the New Yorker Caption Contest and Cartoonstock.com, with expert-annotated evaluation rubrics identifying essential joke elements. LLMs are evaluated based on their explanations towards the humor and abilities in identifying the joke elements. To perform well on this task, models must form and test hypotheses about associations between concepts, potentially backtracking from initial interpretations to arrive at the most plausible explanation. Our extensive benchmarking of current SOTA models reveals three key insights: (1) LLM progress on STEM reasoning transfers effectively to humor comprehension; (2) models trained exclusively on STEM reasoning data still perform well on HumorBench, demonstrating strong transferability of reasoning abilities; and (3) test-time scaling by increasing thinking token budgets yields mixed results across different models in humor reasoning.

幽默理解推理能力大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。