arXiv:2409.13740cs.CLcs.AI2024-09被引 159

语言模型代理在科研文献任务中超越人类专家,准确率更高。

Language agents achieve superhuman synthesis of scientific knowledge

论文配图:Language agents achieve superhuman synthesis of scientific knowledge
图 1 · 摘自论文原文
  • 用真实文献检索任务测试模型,优化事实性
  • 每篇论文发现2.34个矛盾,70%经专家验证
  • 适合科研人员快速查证与综述写作

语言模型常产生错误信息,其在科学研究中的准确性尚不明确。我们设计了一套严格的真人-AI对比方法,评估语言模型代理在真实世界文献搜索任务中的表现,涵盖信息检索、摘要生成和矛盾检测。结果显示,针对事实性优化的前沿模型代理PaperQA2,在无需限制的人类条件下(可全网搜索、使用工具、充足时间),在三项真实科研任务上表现匹配或超越领域专家。PaperQA2撰写的带引用、类似维基百科的科学主题摘要,显著优于现有真人撰写版本。我们还推出了一个硬基准数据集LitQA2,指导PaperQA2的设计,使其性能超过人类。最后,我们将PaperQA2应用于识别科学文献中的矛盾,该任务对人类极具挑战。在随机选取的生物学论文子集中,它平均每篇发现2.34 ± 1.99处矛盾,其中70%被人类专家验证。结果表明,语言模型代理已能在有意义的科研任务中超越领域专家。

原文摘要 · Abstract (English)

Language models are known to hallucinate incorrect information, and it is unclear if they are sufficiently accurate and reliable for use in scientific research. We developed a rigorous human-AI comparison methodology to evaluate language model agents on real-world literature search tasks covering information retrieval, summarization, and contradiction detection tasks. We show that PaperQA2, a frontier language model agent optimized for improved factuality, matches or exceeds subject matter expert performance on three realistic literature research tasks without any restrictions on humans (i.e., full access to internet, search tools, and time). PaperQA2 writes cited, Wikipedia-style summaries of scientific topics that are significantly more accurate than existing, human-written Wikipedia articles. We also introduce a hard benchmark for scientific literature research called LitQA2 that guided design of PaperQA2, leading to it exceeding human performance. Finally, we apply PaperQA2 to identify contradictions within the scientific literature, an important scientific task that is challenging for humans. PaperQA2 identifies 2.34 +/- 1.99 contradictions per paper in a random subset of biology papers, of which 70% are validated by human experts. These results demonstrate that language model agents are now capable of exceeding domain experts across meaningful tasks on scientific literature.

语言模型科研辅助文献挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。