动态评测大模型发现生物新知识的能力,避免数据泄露。
Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
- 构建三阶段自动化流程,从权威论文中生成可测的科学问题与答案。
- 每月更新覆盖12个生物医学领域,确保评估内容新鲜且无训练污染。
- 首次实现动态、自动化的知识发现能力评测,适合研究AI认知能力的学者。
大型语言模型代理在自动知识发现方面展现出巨大潜力,但严格评估其知识发现能力仍具挑战。现有基准多依赖静态数据集,导致模型在训练中可能已接触过测试知识,引发数据污染。同时,现代大模型迭代迅速,静态基准很快过时,无法衡量真正的新知识发现能力。为此,我们提出DBench-Bio,一个动态且全自动的生物知识发现评测框架。该框架包含三个阶段:(1)获取严谨权威的论文摘要;(2)利用大模型生成科学假设问题及其对应发现答案;(3)通过相关性、清晰度和核心性标准筛选高质量问答对。我们基于该流程构建了一个每月更新的基准,覆盖12个生物医学子领域。对当前主流模型的广泛评估揭示了其在发现新知识方面的显著局限。本工作首次提供动态、自动化的评估框架,为人工智能研究社区建立了一个持续演进的知识发现评测资源,推动相关技术发展。
原文摘要 · Abstract (English)
Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rigorously evaluating an AI's capacity for knowledge discovery remains a critical challenge. Existing benchmarks predominantly rely on static datasets, leading to inevitable data contamination where models have likely seen the evaluation knowledge during training. Furthermore, the rapid release cycles of modern LLMs render static benchmarks quickly outdated, failing to assess the ability to discover truly new knowledge. To address these limitations, we propose DBench-Bio, a dynamic and fully automated benchmark designed to evaluate AI's biological knowledge discovery ability. DBench-Bio employs a three-stage pipeline: (1) data acquisition of rigorous, authoritative paper abstracts; (2) QA extraction utilizing LLMs to synthesize scientific hypothesis questions and corresponding discovery answers; and (3) QA filter to ensure quality based on relevance, clarity, and centrality. We instantiate this pipeline to construct a monthly-updated benchmark covering 12 biomedical sub-domains. Extensive evaluations of SOTA models reveal current limitations in discovering new knowledge. Our work provides the first dynamic, automatic framework for assessing the new knowledge discovery capabilities of AI systems, establishing a living, evolving resource for AI research community to catalyze the development of knowledge discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。