用幽默理论指导多阶段推理,让AI生成更懂笑点的图文描述。
HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
- 基于幽默心理学理论构建多阶段推理链,分步解析图像并生成笑点。
- 在多个数据集上超越顶尖模型,人类评分和幽默偏好显著提升。
- 适合研究可解释性生成、幽默理解与跨模态创作的学者与开发者。
幽默既是创造性活动也是社交联结机制,长期挑战AI生成能力。尽管幽默需复杂认知与社会理解,但幽默理论表明其存在可学习的模式与结构,使生成模型可隐式习得。近年来,多模态幽默(如梗图)成为年轻一代主流表达方式,亟需能融合视觉理解与幽默语言生成的AI系统。然而现有数据驱动方法缺乏显式幽默建模或理论基础,常产出流畅但无真正笑点的描述。为此,我们提出HUMORCHAIN(HUmor-guided Multi-step Orchestrated Reasoning Chain for Image Captioning),一个基于幽默与心理学理论的多阶段推理框架。它包含视觉语义解析、幽默与心理驱动推理、微调判别器评估幽默度,形成可解释、可控的认知推理链。据我们所知,这是首个将幽默理论中的认知结构显式嵌入多模态幽默生成的工作,实现从视觉理解到幽默创造的结构化推理。在Meme-Image-No-Text、Oogiri-GO和OxfordTVG-HIC数据集上的实验表明,HUMORCHAIN在人类幽默偏好、Elo/BT评分与语义多样性上均优于当前最优基线,证明理论引导的结构化推理能使大模型生成符合人类感知的幽默内容。
原文摘要 · Abstract (English)
Humor, as both a creative human activity and a social binding mechanism, has long posed a major challenge for AI generation. Although producing humor requires complex cognitive reasoning and social understanding, theories of humor suggest that it follows learnable patterns and structures, making it theoretically possible for generative models to acquire them implicitly. In recent years, multimodal humor has become a prevalent form of online communication, especially among Gen Z, highlighting the need for AI systems capable of integrating visual understanding with humorous language generation. However, existing data-driven approaches lack explicit modeling or theoretical grounding of humor, often producing literal descriptions that fail to capture its underlying cognitive mechanisms, resulting in the generated image descriptions that are fluent but lack genuine humor or cognitive depth. To address this limitation, we propose HUMORCHAIN (HUmor-guided Multi-step Orchestrated Reasoning Chain for Image Captioning), a theory-guided multi-stage reasoning framework. It integrates visual semantic parsing, humor- and psychology-based reasoning, and a fine-tuned discriminator for humor evaluation, forming an interpretable and controllable cognitive reasoning chain. To the best of our knowledge, this is the first work to explicitly embed cognitive structures from humor theories into multimodal humor generation, enabling a structured reasoning process from visual understanding to humor creation. Experiments on Meme-Image-No-Text, Oogiri-GO, and OxfordTVG-HIC datasets show that HUMORCHAIN outperforms state-of-the-art baselines in human humor preference, Elo/BT scores, and semantic diversity, demonstrating that theory-driven structured reasoning enables large language models to generate humor aligned with human perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。