arXiv:2503.02972cs.CLcs.AI2025-03被引 4

用文字伪装技术测试大模型真推理能力,避免靠记忆凑答案。

LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

  • 对语言奥赛题做文字层面的刻意混淆,保留逻辑但难靠记忆解题。
  • 顶尖模型在原题上得分0.59,混淆后降到0.48,仍明显依赖知识。
  • 适合评估模型是否真会推理,尤其关注去知识化的真实思维能力。

前沿语言模型在解决推理问题时表现日益出色,但其性能常因绕过推理过程,转而依赖不断扩展的知识库与记忆能力而被夸大。我们提出 LINGOLY-TOO,一个包含1,203个问题和总计6,995个子问题的挑战性推理基准,通过专家设计的正字学混淆手段作用于语言奥赛题目。这些混淆在保持底层解题逻辑不变的前提下,显著降低题目通过知识或记忆可解的可能性。实验表明,模型在原始问题上利用捷径,而经混淆后性能明显下降。即使最佳推理模型也表现出高度敏感性,得分从原始的约0.59降至0.48。LINGOLY-TOO成功将推理与知识解耦,提供了更清晰的真正推理能力度量。

原文摘要 · Abstract (English)

Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on their expanding knowledge and memorisation capacity. We introduce LINGOLY-TOO, a challenging reasoning benchmark of 1,203 questions and a total of 6,995 sub-questions that counters these shortcuts by applying expert-designed obfuscations to Linguistics Olympiad problems. These obfuscations preserve the underlying solution logic while reducing the likelihood problems are solvable with via knowledge or memorisation. Our experiments show that models exploit shortcuts on the original question as performance markedly drop upon obfuscation. Even the best reasoning models remain highly sensitive, with scores dropping from around 0.59 on original problems to 0.48 after obfuscation. LINGOLY-TOO disentangles reasoning from knowledge, offering a clearer measure of true reasoning capabilities.

推理评测知识解耦语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。