评测大模型在生物医学生成检索中的幻觉问题,揭示其可信度短板。
Overview of TREC 2025 Biomedical Generative Retrieval (BioGen) Track
- 构建TREC 2025生物医学生成检索赛道,聚焦生成内容的真实性评估
- 发现主流大模型在生物医学问答中存在显著幻觉,尤其在关键事实生成上可靠性不足
- 适合关注AI医疗可信性、模型评估与可解释性的研究者参考
大型语言模型(LLMs)在生物医学问答、通俗化文献摘要及临床笔记摘要等任务中取得显著进展,展现出处理和整合复杂生物医学信息并生成流畅自然回复的能力。然而,在高风险领域如医疗问答、临床决策和科研评估中,模型产生的幻觉或虚构内容仍是关键挑战。已有研究表明,这些模型在将生成陈述与可验证来源对齐方面表现不佳,尤其在需要精准事实支持的场景下,其可靠性亟待提升。为此,TREC 2025生物医学生成检索(BioGen)赛道旨在系统评估大模型在生成过程中对真实证据的依赖程度,推动更可信的生物医学生成系统发展。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have made significant progress across multiple biomedical tasks, including biomedical question answering, lay-language summarization of the biomedical literature, and clinical note summarization. These models have demonstrated strong capabilities in processing and synthesizing complex biomedical information and in generating fluent, human-like responses. Despite these advancements, hallucinations or confabulations remain key challenges when using LLMs in biomedical and other high-stakes domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs' abilities to ground generated statements in verifiable sources have shown that models perform significantly
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。