首个面向中文医疗大模型生成内容的可信度验证数据集。
MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- 构建1321个问题、7409条声明的医疗事实核查数据集。
- 发现当前大模型在医学事实核查中存在显著错误率。
- 适合医疗AI安全与可靠性研究者使用。
随着越来越多的人在线获取医疗信息,医学事实核查变得日益重要。然而,现有数据集主要聚焦于人工生成内容,对大型语言模型(LLMs)生成内容的验证仍属空白。为此,我们提出了MedFact,首个基于证据的中文医疗大模型生成内容事实核查数据集。该数据集包含1,321个问题和7,409条陈述,模拟真实医疗场景的复杂性。我们在上下文学习(ICL)和微调两种设置下进行了全面实验,展示了当前大模型在此任务上的能力与挑战,并通过深入的错误分析指明了未来研究的关键方向。数据集已公开发布于https://github.com/AshleyChenNLP/MedFact。
原文摘要 · Abstract (English)
Medical fact-checking has become increasingly critical as more individuals seek medical information online. However, existing datasets predominantly focus on human-generated content, leaving the verification of content generated by large language models (LLMs) relatively unexplored. To address this gap, we introduce MedFact, the first evidence-based Chinese medical fact-checking dataset of LLM-generated medical content. It consists of 1,321 questions and 7,409 claims, mirroring the complexities of real-world medical scenarios. We conduct comprehensive experiments in both in-context learning (ICL) and fine-tuning settings, showcasing the capability and challenges of current LLMs on this task, accompanied by an in-depth error analysis to point out key directions for future research. Our dataset is publicly available at https://github.com/AshleyChenNLP/MedFact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。