arXiv:2412.01621cs.CLcs.AI2024-12被引 8

用纽约时报谜题测试大模型的深度推理能力,发现顶尖模型仍远逊人类。

NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers

  • 设计358个文字分类谜题,阻断直觉思维,考察深层推理。
  • 顶级模型GPT-4准确率比人类低近30%,多轮尝试也难追上。
  • 适合评估大模型真实推理力,尤其对提示工程有效性提出挑战。

大型语言模型在多项基准测试中表现优异,但其进行深思熟虑推理的能力仍存疑。我们提出NYT-Connections,一个由《纽约时报》连接游戏衍生的358个简单词语分类谜题数据集。该基准旨在惩罚快速、直觉式的“系统1”思维,专注于考察基础推理能力。我们在三种配置下评估了六种近期大模型、一种简单机器学习启发式方法及人类表现:单次尝试、多次尝试无提示、多次尝试带上下文提示。结果表明存在显著性能差距:即使表现最佳的GPT-4,准确率也比人类低近30%。值得注意的是,链式思考(Chain-of-Thought)与自洽性(Self-Consistency)等先进提示技术在任务难度升高时收益递减。NYT-Connections结合语言隔离、抵制直觉捷径与定期更新机制,有效防止数据泄露,为评估大模型推理能力提供了新工具。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Connections, a collection of 358 simple word classification puzzles derived from the New York Times Connections game. This benchmark is designed to penalize quick, intuitive "System 1" thinking, isolating fundamental reasoning skills. We evaluated six recent LLMs, a simple machine learning heuristic, and humans across three configurations: single-attempt, multiple attempts without hints, and multiple attempts with contextual hints. Our findings reveal a significant performance gap: even top-performing LLMs like GPT-4 fall short of human performance by nearly 30%. Notably, advanced prompting techniques such as Chain-of-Thought and Self-Consistency show diminishing returns as task difficulty increases. NYT-Connections uniquely combines linguistic isolation, resistance to intuitive shortcuts, and regular updates to mitigate data leakage, offering a novel tool for assessing LLM reasoning capabilities.

大模型评测推理能力提示工程文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。