构建临床遗传学文献推理基准,评估大模型精准理解科研论文的能力。
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- 基于专家标注的ClinGen数据构建真实研究场景的评测集
- 模型在细粒度指令下表现差,推理型模型擅长精细任务但易幻觉
- 适合关注AI辅助临床遗传研究的医学与计算交叉领域研究者
变异和基因解释是个性化医疗与转化生物医学的基础,但传统方法依赖人工且耗时。生成式语言模型(LMs)可加速这一过程,推动基础研究向临床可操作洞察转化。现有基准多聚焦狭窄任务,难以反映真实研究需求。为此,我们提出CGBench,一个基于ClinGen(临床遗传学专家标注文献解读资源)构建的稳健基准,用于评估语言模型在科学文献中的推理能力。CGBench测试模型在三方面的能力:1)按精确规程提取相关实验结果;2)判断证据强度;3)分类并描述实验结果。我们测试了8种不同语言模型,发现尽管模型展现潜力,但在细粒度指令下仍存在显著差距。推理型模型在细粒度任务中表现更优,而非推理型模型在高层级解释上更佳。最后,通过语言模型判别器对比模型解释与人类解释,发现模型即使正确分类证据,仍常出现幻觉或误读结果。CGBench揭示了语言模型在精确解读科学文献中的优势与局限,为临床遗传学及更广泛科学领域的AI研究开辟新方向。
原文摘要 · Abstract (English)
Variant and gene interpretation are fundamental to personalized medicine and translational biomedicine. However, traditional approaches are manual and labor-intensive. Generative language models (LMs) can facilitate this process, accelerating the translation of fundamental research into clinically-actionable insights. While existing benchmarks have attempted to quantify the capabilities of LMs for interpreting scientific data, these studies focus on narrow tasks that do not translate to real-world research. To meet these challenges, we introduce CGBench, a robust benchmark that tests reasoning capabilities of LMs on scientific publications. CGBench is built from ClinGen, a resource of expert-curated literature interpretations in clinical genetics. CGBench measures the ability to 1) extract relevant experimental results following precise protocols and guidelines, 2) judge the strength of evidence, and 3) categorize and describe the relevant outcome of experiments. We test 8 different LMs and find that while models show promise, substantial gaps exist in literature interpretation, especially on fine-grained instructions. Reasoning models excel in fine-grained tasks but non-reasoning models are better at high-level interpretations. Finally, we measure LM explanations against human explanations with an LM judge approach, revealing that models often hallucinate or misinterpret results even when correctly classifying evidence. CGBench reveals strengths and weaknesses of LMs for precise interpretation of scientific publications, opening avenues for future research in AI for clinical genetics and science more broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。