arXiv:2506.01819cs.CLcs.CY2025-06被引 5

评测大模型对职场幽默的判断能力,发现普遍不准确。

Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor

  • 构建职场幽默语料库,标注恰当性特征。
  • 五款大模型对幽默恰当性判断准确率普遍偏低。
  • 适合关注AI价值观对齐与职场应用的研究者。

随着人工智能和大语言模型(LLMs)的进展,自动化日常任务如自动写作日益受到关注。现有研究多聚焦于使大模型与人类价值观对齐,但职场中常见的专业幽默却长期被忽视。为此,我们构建了一个包含职场幽默语句的数据集,并标注了决定其恰当性的特征。对五款主流大语言模型的评估表明,这些模型在判断幽默恰当性方面表现不佳,常出现误判。该研究揭示了当前大模型在理解社会情境化幽默方面的显著局限,呼吁未来工作更重视这一维度。

原文摘要 · Abstract (English)

With the recent advances in Artificial Intelligence (AI) and Large Language Models (LLMs), the automation of daily tasks, like automatic writing, is getting more and more attention. Hence, efforts have focused on aligning LLMs with human values, yet humor, particularly professional industrial humor used in workplaces, has been largely neglected. To address this, we develop a dataset of professional humor statements along with features that determine the appropriateness of each statement. Our evaluation of five LLMs shows that LLMs often struggle to judge the appropriateness of humor accurately.

大模型职场幽默评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。