arXiv:2607.26397cs.CLcs.LG2026-07

提出无需训练的诊断基准,揭示大模型酶分类失败原因。

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

论文配图:Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
图 1 · 摘自论文原文
  • 拆解酶分类为四个独立可测维度:输出结构、外部知识、推理结构、鲁棒性
  • 开放书本访问使闭卷表现大幅提升,证明知识先于推理至关重要
  • 推理实际是权衡冲突邻居,而非获取新知识,适合关注模型决策机制的研究者

酶功能预测是层级化、知识密集型的蛋白质功能分类任务。现有基准存在异常现象:通用大模型在粗粒度第一级分类上表现良好,但要求完整EC编号时,第二至第四级准确率几乎降至零,而专用模型与工具仍保持可用性。本文提出EC-Reason-Bench,一个无需训练的诊断评估协议,旨在回答两个问题:为何通用大模型在EC编号预测上表现近乎为零,以及在不更新任何权重的前提下,能恢复多少性能。我们将酶分类能力分解为四个正交维度:输出结构、外部知识、推理结构和推理鲁棒性,并分别以推理时方法测试各维度表现,对比共享零样本基线(重现此前近零性能)。对多个强推理大模型的实验得出四项主要发现:第一,外部知识起决定性作用,必须先于推理;闭卷表现普遍低下,开放书本后显著提升,缩小了模型差距;第二,在闭卷设置下,级联与思维链的增益或损害取决于模型是否倾向回避;第三,一旦证据可用,最优大模型设置的综合得分与仅投票最近邻检索到的EC编号无异;此差异实为平均效应掩盖,隐藏着在对抗性证据集上大幅增益与多功酶上同等损失;因此推理实质是仲裁冲突邻居,而非知识来源,单一指标排行榜无法捕捉这一本质;第四,准确率服从同源性可用性定律。

原文摘要 · Abstract (English)

Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model's tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.

酶分类大模型诊断推理机制知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。