让AI像漫画编剧一样理解幽默,通过推理过程监督提升表现
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

- 设计三步推理框架:识别画面矛盾、重构合理解释、匹配人类偏好
- 在720亿参数模型上接近专家水平,在纽约客漫画竞赛中领先基线
- 适合研究幽默理解、可解释AI与多模态推理的学者与开发者
幽默是少数既需正确答案又需正确推理的认知任务。尽管现有工作在《纽约客》漫画标题竞赛(NYCC)等基准上评估幽默理解,但大多将其视为黑箱预测,忽视了幽默理解背后的结构化推理过程。本文提出IRS(不一致-化解监督)框架,将幽默理解分解为三个部分:不一致建模(识别视觉场景中的不匹配)、化解建模(构建这些不匹配的连贯重解)和偏好对齐(根据人类判断评估候选解释)。该框架基于不一致-化解理论和专业编剧实践,通过结构化推理轨迹监督中间过程,使从视觉感知到幽默解释的路径清晰可学。在7B、32B和72B规模模型上,IRS在标题匹配与排序任务中均超越强大多模态基线,最大模型在排序任务上接近专家水平。零样本迁移至外部基准显示,IRS学习到了可泛化的推理模式。结果表明,对推理结构的监督比单纯扩大模型规模更为关键。
原文摘要 · Abstract (English)
Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting the answer right. While recent work evaluates humor understanding on benchmarks such as the New Yorker Cartoon Caption Contest (NYCC), it largely treats it as black-box prediction, overlooking the structured reasoning processes underlying humor comprehension. We introduce IRS (Incongruity-Resolution Supervision), a framework that decomposes humor understanding into three components: incongruity modeling, which identifies mismatches in the visual scene; resolution modeling, which constructs coherent reinterpretations of these mismatches; and preference alignment, which evaluates candidate interpretations under human judgments. Grounded in incongruity-resolution theory and expert captionist practice, IRS supervises intermediate reasoning process through structured traces that make the path from visual perception to humorous interpretation explicit and learnable. Across 7B, 32B, and 72B models on NYCC, IRS outperforms strong open and closed multimodal baselines across caption matching and ranking tasks, with our largest model approaching expert-level performance on ranking. Zero-shot transfer to external benchmarks shows that IRS learns generalizable reasoning patterns. Our results suggest that supervising reasoning structure, rather than scale alone, is key for reasoning-centric tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。