评估大模型科学知识更新能力,发现现有方法仍存明显短板
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
- 设计三维度评估框架:保留、获取、预测科学知识
- 最优方法仅保留85.9%旧知识,获取71.7%新知识,预测37.7%未来知识
- 适用于评估科学领域大模型更新机制,对科研辅助工具研发有指导意义
大语言模型在科研中应用日益广泛,但其科学知识易过时。我们提出ScienceMeter框架,用于评估模型在历史、当前及未来科学知识上的更新能力。该框架定义三项指标:知识保留(模型对已有论文理解的保持程度)、知识获取(对新引入论文主张的掌握能力)和知识投影(对潜在未来科学主张的预见与泛化能力)。通过在十个领域的精选数据集上进行命题判断与生成任务,我们评估了五种代表性知识更新方法。结果显示,表现最佳的方法仅能保留85.9%的原有知识,获取71.7%的新知识,对未来知识的投影能力为37.7%,表明构建稳健的科学知识更新机制仍具挑战性且至关重要。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to support scientific research, but their knowledge of scientific advancements can quickly become outdated. We introduce ScienceMeter, a new framework for evaluating scientific knowledge update methods over scientific knowledge spanning the past, present, and future. ScienceMeter defines three metrics: knowledge preservation, the extent to which models' understanding of previously learned papers is preserved; knowledge acquisition, how well scientific claims from newly introduced papers are acquired; and knowledge projection, the ability of the updated model to anticipate or generalize to related scientific claims that may emerge in the future. Using ScienceMeter, we evaluate the scientific knowledge of LLMs through claim judgment and generation tasks on a curated dataset across ten domains. We evaluate five representative knowledge update approaches and find that the best-performing knowledge update methods can preserve only 85.9% of existing knowledge, acquire 71.7% of new knowledge, and project 37.7% of future knowledge, underscoring that developing robust scientific knowledge update mechanisms is both crucial and challenging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。