测试大模型能否选对漫画段子回复,发现它懂幽默但用不好图像。
Memes-as-Replies: Can Models Select Humorous Manga Panel Responses?
- 构建10万组漫画段子与社交帖子配对数据集,评估模型选梗能力。
- 模型能识别夸张等社会线索,但图像信息未提升表现。
- 在细微幽默差异上不如人类,适合研究人机互动幽默理解。
表情包是现代网络交流的常见元素,不仅作为静态内容,也常用于对话中作为互动回应。尽管计算研究关注表情包的内在属性,但其动态、语境化的幽默使用仍是网络科学中的未充分探索领域。为此,我们提出「表情包回复选择」任务,并构建了MaMe-Re(漫画表情包回复基准)数据集,包含10万组由2,325名不同标注者提供的开源日本漫画分镜与社交媒体帖子配对(共50万条标注)。分析显示:(1) 大语言模型初步展现出捕捉夸张等复杂社会线索的能力,超越表面语义匹配;(2) 引入视觉信息并未提升性能,揭示模型理解视觉内容与有效利用其进行语境幽默之间存在差距;(3) 尽管在受控环境中可匹配人类判断,模型仍难以区分语义相近候选回复间的微妙幽默差异。这些发现表明,当前模型在选择情境化幽默回复方面仍面临挑战。
原文摘要 · Abstract (English)
Memes are a popular element of modern web communication, used not only as static artifacts but also as interactive replies within conversations. While computational research has focused on analyzing the intrinsic properties of memes, the dynamic and contextual use of memes to create humor remains an understudied area of web science. To address this gap, we introduce the Meme Reply Selection task and present MaMe-Re (Manga Meme Reply Benchmark), a benchmark of 100,000 human-annotated pairs (500,000 total annotations from 2,325 unique annotators) consisting of openly licensed Japanese manga panels and social media posts. Our analysis reveals three key insights: (1) large language models (LLMs) show preliminary evidence of capturing complex social cues such as exaggeration, moving beyond surface-level semantic matching; (2) the inclusion of visual information does not improve performance, revealing a gap between understanding visual content and effectively using it for contextual humor; (3) while LLMs can match human judgments in controlled settings, they struggle to distinguish subtle differences in wit among semantically similar candidates. These findings suggest that selecting contextually humorous replies remains an open challenge for current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。