arXiv:2503.13517cs.CLcs.AI2025-03ICLR被引 41

测试大模型在科学长文本理解与推理能力,发现现有模型仍有巨大提升空间。

CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

  • 构建跨六学科的科学长文本任务集,涵盖实验与理论工作流。
  • 顶尖模型最高仅32%正确率,蛋白质序列任务表现最差。
  • 适合关注AI辅助科研、模型评估的研究者参考。

科学问题解决需要整合信息并运用专业知识。我们提出CURIE,一个用于评估大语言模型(LLMs)在科学长文本理解、推理和信息提取方面潜力的基准。该基准包含由六个领域专家精心设计的10个挑战性任务,共580个问题及解答对,覆盖材料科学、凝聚态物理、量子计算、地理空间分析、生物多样性和蛋白质研究,涉及科学中的实验与理论工作流程。我们在CURIE上评估了多种封闭与开源的LLMs,任务要求具备领域知识、长上下文理解能力及多步推理。尽管Gemini Flash 2.0和Claude-3在各领域表现出一致的高理解能力,但主流GPT-4o和command-R+在蛋白质序列任务中表现严重失败。最佳模型仅达32%准确率,表明所有模型均有巨大提升空间。我们希望CURIE带来的洞见能推动未来科学领域LLM的发展。评估代码与数据见https://github.com/google/curie。

原文摘要 · Abstract (English)

Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding,Reasoning and Information Extraction benchmark to measure the potential of Large Language Models (LLMs) in scientific problem-solving and assisting scientists in realistic workflows. This benchmark introduces ten challenging tasks with a total of 580 problems and solution pairs curated by experts in six disciplines - materials science, condensed matter physics, quantum computing, geospatial analysis, biodiversity, and proteins - covering both experimental and theoretical work-flows in science. We evaluate a range of closed and open LLMs on tasks in CURIE which requires domain expertise, comprehension of long in-context information,and multi-step reasoning. While Gemini Flash 2.0 and Claude-3 show consistent high comprehension across domains, the popular GPT-4o and command-R+ fail dramatically on protein sequencing tasks. With the best performance at 32% there is much room for improvement for all models. We hope that insights gained from CURIE can guide the future development of LLMs in sciences. Evaluation code and data are in https://github.com/google/curie

大模型评测科学智能长文本理解多任务推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。