arXiv:2410.22394cs.CL2024-10ICML被引 27

评测大模型在科研任务中的表现,发现其潜力与局限。

AAAR-1.0: Assessing AI's Potential to Assist Research

  • 构建面向科研的四类任务评估集,需深度专业理解。
  • 大模型在推导方程、设计实验等任务上表现尚可但不完善。
  • 适合关注AI辅助科研的学者和研究工具开发者。

大量研究已评估大语言模型(LLMs)在日常任务如写邮件、问答和创意生成中的能力,但研究人员在构思研究思路、设计实验、撰写或评审论文时面临独特挑战。本文提出AAAR-1.0,一个专为评估LLM在四项高阶科研任务中表现而设计的基准数据集:(i) EquationInference,基于论文上下文判断方程正确性;(ii) ExperimentDesign,设计实验验证研究想法;(iii) PaperWeakness,识别论文提交中的弱点;(iv) REVIEWCRITIQUE,判断人工评审各段落的缺陷。该数据集区别于以往基准在于:一是明确聚焦科研场景,任务需深层领域知识;二是贴近研究人员日常实践。对开源与专有模型的评估揭示了其在复杂科研任务中的潜力与不足。我们将持续迭代更新至后续版本。

原文摘要 · Abstract (English)

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face unique challenges and opportunities in leveraging LLMs for their own work, such as brainstorming research ideas, designing experiments, and writing or reviewing papers. In this study, we introduce AAAR-1.0, a benchmark dataset designed to evaluate LLM performance in three fundamental, expertise-intensive research tasks: (i) EquationInference, assessing the correctness of equations based on the contextual information in paper submissions; (ii) ExperimentDesign, designing experiments to validate research ideas and solutions; (iii) PaperWeakness, identifying weaknesses in paper submissions; and (iv) REVIEWCRITIQUE, identifying each segment in human reviews is deficient or not. AAAR-1.0 differs from prior benchmarks in two key ways: first, it is explicitly research-oriented, with tasks requiring deep domain expertise; second, it is researcher-oriented, mirroring the primary activities that researchers engage in on a daily basis. An evaluation of both open-source and proprietary LLMs reveals their potential as well as limitations in conducting sophisticated research tasks. We will keep iterating AAAR-1.0 to new versions.

AI科研大模型评估研究助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。