测试大模型理解与创作中文歇后语的能力,发现其表现依赖记忆而非真正推理。
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

- 用新创歇后语测试模型,通过准确率差异判断是否依赖记忆。
- 顶尖中文模型在新歇后语上准确率差达23.6%,远超英文模型的5.1%。
- 模型生成的歇后语质量远低于人类,说明其创造性仍不足。
本文通过测试大语言模型在中文歇后语这一语言游戏中的表现,推进了对模型推理能力的边界探索。我们使用语言学家新创的、此前未出现过的歇后语以避免数据泄露问题,采用多选题(MCQ)、自由解释生成和新歇后语创造三种任务评估模型的理解与生成能力。在多选题中,以现有但低频歇后语与新歇后语之间的准确率差异(Δ_acc)作为记忆指标:母语者Δ_acc极低,表明处理机制相似;而前沿中文模型平均Δ_acc达23.6%,英语主导模型则为5.1%,表明中文模型更可能因训练数据量大而记忆更多低频内容。在新歇后语生成任务中,Gemini 3.1 Pro 准确率达92.6%,比人类高24%;但在歇后语创作任务中,模型生成结果评分远低于人类。这些结果提示,在存在数据污染风险的情况下,对大模型推理能力的宣称需谨慎审视,且其在语言创作类任务中,尤其在中文歇后语领域,仍不及人类专家。
原文摘要 · Abstract (English)
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($Δ_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $Δ_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $Δ_{acc}$ of 23.6\%, while English-centric models tested have a mean $Δ_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。