arXiv:2605.13408cs.CL2026-05

构建语言谜题配对数据集,对比人类与大模型解题表现

From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks

  • 将罗塞塔石碑谜题自动转换为匹配题,提升出题效率
  • 人类和大模型在匹配题上呈现全对或全错的二元解题模式
  • 适合研究语言推理能力、人机对比或竞赛题生成的学者

本文研究高中语言竞赛中常见的两类语言谜题:罗塞塔石碑题与匹配题。提出一种系统性方法,可将现有罗塞塔石碑题自动转换为对应的匹配题,显著提升新谜题生成效率。通过人类参与者和大语言模型(LLMs)评估生成的配对谜题,结果表明,无论是专家级人类解题者还是大模型,在匹配题上均表现出‘全对或全错’的二元模式。本工作贡献了一个配对谜题数据集,并详细分析了不同题型下的难度差异,为理解人类与机器的语言推理能力提供了新视角。

原文摘要 · Abstract (English)

In this paper, we examine linguistic puzzles used in high school linguistics competitions, focusing on two common formats: Rosetta Stone and Match-Up. We propose a systematic procedure for converting existing Rosetta Stone puzzles into corresponding Match-Up counterparts. Because linguistic puzzle creation is complex and time-consuming, our method provides an efficient way to accelerate the generation of new puzzles. We evaluate the resulting Rosetta Stone-Match-Up pairs with both human participants and large language models (LLMs). Our results show that both expert human solvers and LLMs display an all-or-nothing pattern on Match-Up puzzles, either solving them completely or failing entirely. This work contributes a new dataset of paired puzzles and provides a detailed evaluation of puzzle difficulty across formats, offering insights into both human and machine linguistic reasoning.

语言推理人机对比谜题生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。