arXiv:2604.08757cs.CLcs.AI2026-04被引 2

用幽默卡牌游戏测试大模型笑点判断能力,发现模型偏好与人类差异显著。

Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models

  • 让五款前沿大模型玩《人类荒诞卡》游戏,从10张卡中选最搞笑的
  • 模型平均表现优于随机水平,但与人类偏好一致度不高
  • 模型间一致性远超人机一致性,可能受位置和内容偏见影响

幽默是人类交流中文化嵌入最深、社会意义最重的维度之一,但在大语言模型对齐研究中仍鲜有探讨。本研究让五款前沿语言模型与人类玩家同玩《人类荒诞卡》(Cards Against Humanity, CAH)游戏,在9,894轮中从十张候选卡中选出最有趣的回应。尽管所有模型均显著优于随机基线,但其与人类偏好的一致性仍较有限。更令人惊讶的是,模型之间的共识程度远高于与人类的共识。我们发现这一现象部分可由系统性位置偏差和内容偏好解释,从而引发疑问:大模型的幽默判断究竟反映真实偏好,还是推理与对齐过程中的结构性产物?

原文摘要 · Abstract (English)

Humor is one of the most culturally embedded and socially significant dimensions of human communication, yet it remains largely unexplored as a dimension of Large Language Model (LLM) alignment. In this study, five frontier language models play the same Cards Against Humanity games (CAH) as human players. The models select the funniest response from a slate of ten candidate cards across 9,894 rounds. While all models exceed the random baseline, alignment with human preference remains modest. More striking is that models agree with each other substantially more often than they agree with humans. We show that this preference is partly explained by systematic position biases and content preferences, raising the question whether LLM humor judgment reflects genuine preference or structural artifacts of inference and alignment.

幽默理解模型对齐卡牌游戏人类偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。