arXiv:2510.10549cs.AI2025-10被引 2

评测大模型读论文能力,发现其理解深度远不如人类。

ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding

  • 专家设计3种难度题目,强调深层推理而非简单检索。
  • 顶尖模型准确率仅39.95%,显著低于人类水平。
  • 思考模式或检索增强反而降低表现,暴露模型缺陷。

尽管大语言模型在诸多领域任务中表现优异,但其对完整学术论文的深度理解与推理能力仍缺乏深入探索。现有基准多因问题设计肤浅或评估指标不可靠而难以反映真实理解水平。为此,我们推出ELAIPBench,一个由领域专家构建的基准,用于评估大模型对人工智能研究论文的理解能力。该基准通过激励性对抗标注流程,涵盖137篇论文中的403道多选题,分为三个难度层级,强调非平凡推理而非表层信息提取。实验表明,表现最佳的模型准确率仅为39.95%,远低于人类水平。此外,具备思维链或检索增强生成(RAG)能力的前沿模型未能提升性能,甚至因过度思考或噪声检索导致准确率下降。这些结果凸显当前大模型在真正理解学术论文方面仍存在巨大差距。

原文摘要 · Abstract (English)

While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Existing benchmarks often fall short of capturing such depth, either due to surface-level question design or unreliable evaluation metrics. To address this gap, we introduce ELAIPBench, a benchmark curated by domain experts to evaluate LLMs' comprehension of artificial intelligence (AI) research papers. Developed through an incentive-driven, adversarial annotation process, ELAIPBench features 403 multiple-choice questions from 137 papers. It spans three difficulty levels and emphasizes non-trivial reasoning rather than shallow retrieval. Our experiments show that the best-performing LLM achieves an accuracy of only 39.95%, far below human performance. Moreover, we observe that frontier LLMs equipped with a thinking mode or a retrieval-augmented generation (RAG) system fail to improve final results-even harming accuracy due to overthinking or noisy retrieval. These findings underscore the significant gap between current LLM capabilities and genuine comprehension of academic papers.

论文理解大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。