新基准测试发现大模型推理能力多依赖知识而非逻辑结构。
IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
- 用同构跨领域题目分离逻辑推理与知识检索能力。
- 91.3%的推理提升实际来自知识依赖,非结构通用性。
- 不同评测集会得出相反结论,选基准很重要。
我们提出ISOSCI,一个同构跨领域科学问题对的基准,用于分离大模型评估中的推理能力与领域知识检索。每对问题具有相同逻辑结构但需不同领域知识,实现推理收益的可控归因。在涵盖四个模型家族的五组模型中,91.3%的推理模式增益为知识依赖型(63/69次增益;Wilson 95%置信区间[82.3%, 96.0%]),直接挑战链式思维推理可提升短程程序化科学问题解决能力的假设。在所有领域中,高能力模型开启推理模式带来的准确率提升均不足5个百分点;而专门优化推理的o3-mini模型在GPQA Diamond上优于标准版本19.2个百分点,但在ISOSCI上却落后24.7个百分点,表明评测基准选择决定了对推理有效性的判断。ISOSCI已发布于https://huggingface.co/datasets/isosci/isosci。
原文摘要 · Abstract (English)
We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation. Each pair shares identical logical structure but requires different domain-specific knowledge, enabling controlled attribution of reasoning-mode gains. Across five model pairs spanning four model families, we find that 91.3% of reasoning-mode gains are knowledge-dependent rather than structure-invariant (63/69 gains; Wilson 95% CI [82.3%, 96.0%]), directly challenging the assumption that chain-of-thought reasoning improves short-horizon procedural scientific problem-solving. Reasoning toggles on highly capable models provide less than 5 percentage points accuracy gain across all domains, and a reasoning-specialized model (o3-mini) that outperforms its standard counterpart on GPQA Diamond (+19.2 percentage points) underperforms on ISOSCI (-24.7 percentage points), showing that benchmark choice determines conclusions about reasoning utility. We release ISOSCI at https://huggingface.co/datasets/isosci/isosci
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。