让幽默生成更懂观众,用人类偏好判断选出最好笑的笑话。
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
- 通过多轮提示和模型集成生成多样笑话候选
- 基于2500组人类对比数据训练偏好模型,精准捕捉观众喜好
- 系统在中英西三语任务中表现领先,支持后续研究复现
幽默生成不仅因创作流畅新颖的笑话困难,更因‘好笑’取决于受众、情境与文化,且标注噪音大、人工一致性低。本文介绍针对SemEval-2026任务1(MWAHAHA)的系统,该任务聚焦受约束的幽默生成,通过一对一竞技式人类偏好判断评估。我们采用“生成多条→选择最佳”策略:首先利用多步提示、模型集成与多样性解码生成多样化候选;其次构建偏好模型,通过学习人类对比而非绝对搞笑分来模拟读者视角。为此,我们发布了2500组通过“幽默竞技场”原型收集的人类成对判断数据。进一步提出可解释流水线,将标注对比转化为偏好模型。在三个偏好数据集上,模型均优于基线并展现更强跨领域迁移能力。最终,我们将学习到的偏好模型用于MWAHAHA设置下的候选排序,并公开中间成果(候选池与排序结果)以促进后续研究。系统在英语和中文子任务中排名第一,在西班牙语子任务中排名第二。
原文摘要 · Abstract (English)
Humor generation remains difficult not only because producing fluent, novel jokes is hard, but because "funny" is audience-dependent and supervision is noisy -- preferences vary with audience, context, and culture, and annotator agreement is often low. In this paper, we describe our system for the SemEval-2026 Task-1 (MWAHAHA), which focuses on humor generation under explicit constraints. The task evaluates submitted systems via human preference judgments in 1-on-1 arena-style comparisons. We adopt a "generate-many -> select-best" strategy. First, we generate a diverse pool of candidates per instance using multi-step prompting, model ensembling, and diversity-oriented decoding. Second, we select outputs using a preference model that approximates a "reader" by learning from human comparisons rather than absolute funniness scores. To support this approach, we release 2.5K human pairwise judgments collected through the Humor Arena prototype. We further propose an interpretable pipeline that converts labeled comparisons into a preference model. Across three preference datasets, our models consistently outperform baselines and show stronger cross-domain transfer. Finally, we apply the learned preference model to rank candidates for the MWAHAHA setting and release intermediate artifacts (candidate pools and rankings) to facilitate follow-up work. Our system ranked 1st in the English and Chinese subtasks of MWAHAHA and 2nd in the Spanish subtask.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。