构建中文网络梗解释数据集,测试大模型对流行梗的理解能力
Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation
- 构建CHIME数据集,包含中文流行语梗的含义、起源等标注
- 大模型能解释部分梗但对文化语境复杂梗表现差,起源识别准确率低
- 适合研究网络文化理解、大模型社会认知能力的学者使用
大型语言模型(LLMs)虽在海量互联网文本上训练,但是否真正理解快速传播的网络流行内容——即“梗”?本文提出CHIME数据集,用于中文网络梗解释。该数据集涵盖中国互联网流行的短语类梗,附有含义、起源、例句、类型等详细标注。我们设计两项任务评估模型理解能力:第一项要求解释梗的含义、识别起源并生成例句,结果显示模型对文化语言复杂的梗理解力显著下降,且难以准确溯源;第二项为多项选择题,需选出最合适的梗填入上下文句子,尽管模型能作答,但表现仍远低于人类水平。我们已公开CHIME数据集,以推动计算式梗理解研究。
原文摘要 · Abstract (English)
Large language models (LLMs) are trained on vast amounts of text from the Internet, but do they truly understand the viral content that rapidly spreads online -- commonly known as memes? In this paper, we introduce CHIME, a dataset for CHinese Internet Meme Explanation. The dataset comprises popular phrase-based memes from the Chinese Internet, annotated with detailed information on their meaning, origin, example sentences, types, etc. To evaluate whether LLMs understand these memes, we designed two tasks. In the first task, we assessed the models' ability to explain a given meme, identify its origin, and generate appropriate example sentences. The results show that while LLMs can explain the meanings of some memes, their performance declines significantly for culturally and linguistically nuanced meme types. Additionally, they consistently struggle to provide accurate origins for the memes. In the second task, we created a set of multiple-choice questions (MCQs) requiring LLMs to select the most appropriate meme to fill in a blank within a contextual sentence. While the evaluated models were able to provide correct answers, their performance remains noticeably below human levels. We have made CHIME public and hope it will facilitate future research on computational meme understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。