arXiv:2604.19786cs.CL2026-04

用比赛排名法评估大模型幽默生成能力,让不同模型可比可解释。

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models

论文配图:HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
图 1 · 摘自论文原文
  • 基于幽默理论设计裁判系统,自动比较模型输出并给出理由。
  • 9个模型在两个基准上排名稳定,小模型也能媲美大模型。
  • 适合研究幽默生成、评测方法或想理解模型幽默机制的人。

评估大语言模型的幽默能力仍是开放挑战,因现有方法产生孤立、不可比的指标,难以追踪系统进展。我们提出HumorRank,一种基于锦标赛的文本幽默生成评估框架与排行榜。在两个公开基准(SemEval-2026 MWAHAHA 和 Humor Transfer Bench)上,对九种模型(涵盖专有、开源和专用系统)进行了广泛的自动化成对评估。裁判由基于通用言语幽默理论(GTVH)的LLM担任,每个裁判结合结构化喜剧分析进行判断,共同输出偏好决策、可解释理由,以及机制、传达和失败标签,而非黑箱趣味分。判断结果通过自适应瑞士轮赛制聚合,采用布拉德利-泰瑞最大似然估计(MLE)生成全局一致的幽默生成能力排名。排名跨裁判稳定:独立使用Llama 3.3 70B和Qwen 2.5 72B的裁判在两个基准上均获得肯德尔τ相关系数0.889;人类校准研究显示,人类与LLM在困难幽默对上的判断一致性接近人类之间的一致性。结果表明,HumorRank能生成统计可靠的模型分层,幽默质量与对矛盾、简洁、升级、荒诞等喜剧机制的掌握有关,而非仅依赖模型规模,专用微调模型可达到远超其规模的大模型水平。HumorRank提供了一种可扩展、可解释、可复现的评测与理解大模型幽默生成的方法。

原文摘要 · Abstract (English)

Evaluating humor in large language models (LLMs) is an open challenge because existing approaches yield isolated, incomparable metrics rather than unified model rankings, making it difficult to track progress across systems. We introduce HumorRank, a tournament-based evaluation framework and leaderboard for textual humor generation. On two public benchmarks (SemEval-2026 MWAHAHA and Humor Transfer Bench), we conduct extensive automated pairwise evaluation across nine models spanning proprietary, open-weight, and specialized systems. Pairwise judgments are produced by LLM judges grounded in the General Theory of Verbal Humor (GTVH): each judge integrates structured comedic analysis into adjudication, jointly yielding a preference decision, an interpretable rationale, and mechanism, delivery, and failure tags rather than a black-box funniness score. Judgments are aggregated via an Adaptive Swiss tournament, with Bradley-Terry Maximum Likelihood Estimation (MLE) producing globally consistent humor generation capability rankings. Rankings are cross-judge stable: independent LLM judges (Llama 3.3 70B and Qwen 2.5 72B) yield Kendall tau = 0.889 on both benchmarks, and a human calibration study shows human-LLM agreement tracking human-human agreement on hard funny-versus-funny pairs. Our results demonstrate that HumorRank yields statistically grounded model stratifications, showing that humor quality is associated with mastery of comedic mechanisms such as incongruity, conciseness, escalation, and absurdity rather than model scale alone, with specialized fine-tuned models reaching parity with far larger systems. HumorRank thus provides a scalable, interpretable, and reproducible methodology for benchmarking and understanding LLM-generated humor.

幽默生成模型评测可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。