用高温超导领域专家问题评估大模型,发现检索增强模型表现更优。
Expert Evaluation of LLM World Models: A High-$T_c$ Superconductivity Case Study
- 构建1726篇论文数据库与67个深度问题,模拟专家阅读理解
- 检索增强模型在事实全面性和证据支持上显著优于闭源模型
- 提出可复用的专家评估标准,适合测试科学推理能力
大型语言模型在科学文献探索中展现出巨大潜力,但其在专业领域回答复杂问题的准确性与全面性仍待研究。以高温铜氧化物超导体为例,我们构建了一个包含1726篇科学论文的专家整理数据库,并设计了67个由专家提出的深度问题,用于检验模型对文献的理解水平。我们评估了六种不同的基于LLM的系统,包括商用闭源模型和一种可检索图文信息的定制检索增强生成(RAG)系统。专家依据平衡视角、事实全面性、简洁性及证据支持等维度对答案进行评分。结果显示,两个采用经专家筛选文献的RAG系统在关键指标上超越现有闭源模型,尤其在提供全面且有依据的回答方面表现突出。本文还讨论了模型的潜力与局限性,并指出所提出的专家问题集与评估框架对未来评估科学推理类大模型具有重要价值。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comprehensive answers to complex questions within specialized domains remains an active area of research. Using the field of high-temperature cuprates as an exemplar, we evaluate the ability of LLM systems to understand the literature at the level of an expert. We construct an expert-curated database of 1,726 scientific papers that covers the history of the field, and a set of 67 expert-formulated questions that probe deep understanding of the literature. We then evaluate six different LLM-based systems for answering these questions, including both commercially available closed models and a custom retrieval-augmented generation (RAG) system capable of retrieving images alongside text. Experts then evaluate the answers of these systems against a rubric that assesses balanced perspectives, factual comprehensiveness, succinctness, and evidentiary support. Among the six systems two using RAG on curated literature outperformed existing closed models across key metrics, particularly in providing comprehensive and well-supported answers. We discuss promising aspects of LLM performances as well as critical short-comings of all the models. The set of expert-formulated questions and the rubric will be valuable for assessing expert level performance of LLM based reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。