模型理解幽默能力与编码器相当,解码器表现媲美顶尖编码器。
Decoders Laugh as Loud as Encoders
- 用微调的GPT-4o解码器评估幽默理解能力
- 解码器在幽默识别上达到0.85的平均F1-macro分数
- 首次证明解码器在幽默任务中可媲美顶级编码器
自计算机诞生以来,艾伦·图灵便梦想着能用语言与人类交流的机器人。近年来,大语言模型(LLMs)的发展令科学界震惊:单一模型可完成多种自然语言处理任务,输出结果甚至优于多数人类沟通能力。GPT、Claude、Grok等模型已在学术界留下印记。然而,这些模型生成内容时的理解深度仍不明确,尤其是在幽默这类细微主题上。目前尚不清楚计算机是否真正理解幽默——此前检查的最新解码器仅为GPT-2。本文通过实证研究解决该问题:我们发现微调后的解码器(GPT-4o)在幽默识别任务中取得0.85的平均F1-macro得分,与最佳微调编码器(RoBERTa)的0.86得分相当。
原文摘要 · Abstract (English)
From the dawn of the computer, Allen Turing dreamed of a robot that could communicate using language as a human being. The recent advances in the field of Large Language Models (LLMs) shocked the scientific community when a single model can apply for various natural language processing (NLP) tasks, while the output results are sometimes even better than most human communication skills. Models such as GPT, Claude, Grok, etc. have left their mark on the scientific community. However, it is unclear how much these models understand what they produce, especially in a nuanced theme such as humor. The question of whether computers understand humor is still open (among the decoders, the latest to be checked was GPT-2). We addressed this issue in this paper; we have showed that a fine-tuned decoder (GPT-4o) performed (Mean F1-macro score of 0.85) as well as the best fine-tuned encoder (RoBERTa with a Mean of F1-score 0.86)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。