arXiv:2607.19011cs.CLcs.AI2026-07

解析图文笑话理解与生成的挑战与进展

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

论文配图:Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
图 1 · 摘自论文原文
  • 按识别、解读、生成三层架构梳理视觉幽默研究
  • 指出当前评估易走捷径,文化覆盖不足
  • 适合关注多模态理解和生成的AI研究者

图文笑话、漫画和连环画中的幽默对人工智能仍具挑战,因其依赖非字面意义、共享文化知识和交际意图,而非字面场景描述。本综述聚焦单图与多格图像中的视觉幽默理解,将幽默生成视为新兴下游任务。在与以往幽默、讽刺及通用多模态大模型综述对比的基础上,构建以能力为中心的层级框架,涵盖识别、解释与推理、生成三阶段。该框架整合了基准设计、评估协议与建模范式,揭示领域从特定任务融合模型向基于多模态对齐、证据驱动推理与可控生成的大模型方法演进。最后,指出现存主要障碍:评估易产生捷径、文化与叙事覆盖有限、证据支撑薄弱,以及安全与版权问题未解。

原文摘要 · Abstract (English)

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

视觉幽默多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。