构建医学幻觉检测数据集,助力AI生成医疗文本更可信
MedHal: An Evaluation Dataset for Medical Hallucination Detection
- 覆盖多种医疗文本任务,标注海量真实病例数据
- 提供事实错误解释,帮助模型学习判断依据
- 适合医疗AI研发者与评测人员快速验证系统可靠性
我们提出MedHal,一个专门用于评估模型在医疗文本中识别幻觉能力的大规模数据集。现有幻觉检测方法在医学等专业领域表现受限,可能带来严重后果。现有医疗数据集规模小(仅数百样本)或仅聚焦单一任务(如问答、自然语言推理)。MedHal通过三方面填补空白:(1) 融合多样化的医疗文本来源与任务;(2) 提供足够数量的标注样本,支持医疗幻觉检测模型训练;(3) 包含事实不一致的详细解释,指导模型学习。我们通过训练并评估基线模型,证明其在医疗幻觉检测上优于通用方法。该资源可高效评估医疗文本生成系统,减少对昂贵专家评审的依赖,有望加速医疗AI研究进展。
原文摘要 · Abstract (English)
We present MedHal, a novel large-scale dataset specifically designed to evaluate if models can detect hallucinations in medical texts. Current hallucination detection methods face significant limitations when applied to specialized domains like medicine, where they can have disastrous consequences. Existing medical datasets are either too small, containing only a few hundred samples, or focus on a single task like Question Answering or Natural Language Inference. MedHal addresses these gaps by: (1) incorporating diverse medical text sources and tasks; (2) providing a substantial volume of annotated samples suitable for training medical hallucination detection models; and (3) including explanations for factual inconsistencies to guide model learning. We demonstrate MedHal's utility by training and evaluating a baseline medical hallucination detection model, showing improvements over general-purpose hallucination detection approaches. This resource enables more efficient evaluation of medical text generation systems while reducing reliance on costly expert review, potentially accelerating the development of medical AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。