TREC 2024设生物医学生成检索任务,解决大模型幻觉问题。
Overview of TREC 2024 Biomedical Generative Retrieval (BioGen) Track
- 引入参考归属任务,强制模型生成时引用可验证来源。
- 发现模型在面向普通用户的提问上引用率显著下降。
- 适合关注医疗AI可信度与可解释性的研究者使用。
随着大语言模型(LLMs)的发展,生物医学领域在问答、通俗化文献摘要、临床笔记摘要等任务上取得显著进展。然而,幻觉或虚构内容仍是使用LLMs于生物医学及其他领域的关键挑战。不准确信息在医疗问答、临床决策或科研评估等高风险场景中可能造成严重后果。已有研究显示,模型在处理普通用户提出的问题时,其生成内容的可追溯性显著下降,常无法引用相关文献。当用户需要证据支持模型结论时,缺乏依据的陈述成为阻碍其应用于健康相关场景的主要障碍。因此,亟需能够将生成内容锚定在可靠来源的方法及实用的评估手段。为此,我们在TREC 2024的试点任务中引入了参考归属(reference attribution)任务,旨在通过要求模型回答时提供可验证的来源,减少虚假陈述的生成。
原文摘要 · Abstract (English)
With the advancement of large language models (LLMs), the biomedical domain has seen significant progress and improvement in multiple tasks such as biomedical question answering, lay language summarization of the biomedical literature, clinical note summarization, etc. However, hallucinations or confabulations remain one of the key challenges when using LLMs in the biomedical and other domains. Inaccuracies may be particularly harmful in high-risk situations, such as medical question answering, making clinical decisions, or appraising biomedical research. Studies on the evaluation of the LLMs abilities to ground generated statements in verifiable sources have shown that models perform significantly worse on lay-user-generated questions, and often fail to reference relevant sources. This can be problematic when those seeking information want evidence from studies to back up the claims from LLMs. Unsupported statements are a major barrier to using LLMs in any applications that may affect health. Methods for grounding generated statements in reliable sources along with practical evaluation approaches are needed to overcome this barrier. Towards this, in our pilot task organized at TREC 2024, we introduced the task of reference attribution as a means to mitigate the generation of false statements by LLMs answering biomedical questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。