arXiv:2501.04249cs.CL2025-01被引 11

用语言奥赛题测试大模型推理能力,发现其规则抽象仍不足。

IOLBENCH: Benchmarking LLMs on Linguistic Reasoning

  • 基于国际语言学奥林匹克竞赛题构建新基准
  • 顶尖模型在构词与句法推理上表现有限
  • 适合研究模型语言逻辑推理的学者参考

尽管深度神经网络取得了显著进展并广泛应用,其进行推理任务的能力仍受限,尤其在需要结构化、抽象思维的领域。本文通过引入IOLBENCH——一个源自国际语言学奥林匹克(IOL)问题的新基准,考察当前最先进的大语言模型(LLMs)在语言推理方面的能力。该数据集涵盖语法、形态、音系和语义等多样题目,均设计为自包含且不依赖外部知识。这些任务要求模型从少量示例中推断语言规则与模式,进行元认知语言推理。通过对主流大模型的广泛评测,我们发现即使最先进模型也难以应对语言复杂性的挑战,特别是在组合泛化与规则抽象方面表现不佳。分析揭示了当前模型在语言问题求解中的优势与持续局限,为理解其推理能力提供了重要洞见。通过引入IOLBENCH,旨在推动更接近人类水平推理模型的发展,对计算语言学与人工智能领域具有广泛影响。

原文摘要 · Abstract (English)

Despite the remarkable advancements and widespread applications of deep neural networks, their ability to perform reasoning tasks remains limited, particularly in domains requiring structured, abstract thought. In this paper, we investigate the linguistic reasoning capabilities of state-of-the-art large language models (LLMs) by introducing IOLBENCH, a novel benchmark derived from International Linguistics Olympiad (IOL) problems. This dataset encompasses diverse problems testing syntax, morphology, phonology, and semantics, all carefully designed to be self-contained and independent of external knowledge. These tasks challenge models to engage in metacognitive linguistic reasoning, requiring the deduction of linguistic rules and patterns from minimal examples. Through extensive benchmarking of leading LLMs, we find that even the most advanced models struggle to handle the intricacies of linguistic complexity, particularly in areas demanding compositional generalization and rule abstraction. Our analysis highlights both the strengths and persistent limitations of current models in linguistic problem-solving, offering valuable insights into their reasoning capabilities. By introducing IOLBENCH, we aim to foster further research into developing models capable of human-like reasoning, with broader implications for the fields of computational linguistics and artificial intelligence.

语言推理大模型评测规则抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。