arXiv:2604.09629cs.CL2026-04被引 1

用心理角色生成幽默数据,让小模型笑得比大模型还妙。

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation

论文配图:HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation
图 1 · 摘自论文原文
  • 设计六种喜剧人格,用思维混合生成多样笑点
  • 70亿参数模型超越更大指令模型,性能达开源顶尖水平
  • 证明幽默关键在数据设计,而非模型大小或对齐算法

幽默生成对大语言模型构成重大挑战,因其标准训练目标(下一个词预测)与喜剧所需的意外和不协调本质冲突。为此,我们提出认知协同框架,基于心理学幽默理论生成高质量幽默数据。采用思维混合(MoT)方法,部署六种认知人格(如荒诞者、犬儒者)为给定提示生成多元喜剧视角。该框架构建出有理论依据的数据集,并用于微调一个70亿参数的学生模型。我们进一步评估了两种对齐策略:直接偏好优化(DPO)和离线组相对变体O-GRPO,发现二者均未优于监督微调(SFT)。然而,我们的70亿参数HumorGen模型显著超越更大的指令微调基线,在开源模型中达到顶尖表现,且与前沿闭源系统相当。结果表明,对于幽默生成而言,认知驱动的数据构建比对齐算法或模型规模更为关键。

原文摘要 · Abstract (English)

Humor generation poses a significant challenge for Large Language Models (LLMs), because their standard training objective (next-token prediction) inherently conflicts with the surprise and incongruity required for comedy. To bridge this gap, we introduce the Cognitive Synergy Framework, a methodology for generating highquality humor data inspired by psychological theories of humor. Utilizing a Mixtureof-Thought (MoT) approach, we deploy six cognitive personas (e.g., The Absurdist, The Cynic) to synthesize diverse comedic perspectives for a given prompt. This framework produces a theory-grounded dataset, which we use to fine-tune a 7B-parameter student model. We further evaluate two alignment strategies, Direct Preference Optimization (DPO) and an offline group-relative variant O-GRPO, finding that neither improves over SFT. However, our 7B HumorGen model variants significantly outperform larger instruction-tuned baselines and achieve top-tier open-weight performance while remaining competitive with frontier proprietary systems. These results suggest that cognitively driven data curation is more critical than alignment algorithms or model scale for humor generation.

幽默生成认知模型数据构建小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。