构建生物医学长文本幻觉评估基准,量化大模型事实错误
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
- 将生成内容拆解为原子信息单元,逐项评分
- 8573个问题覆盖29类,发现模型超70亿参数后效果趋于稳定
- 适合医疗AI评测、模型优化与跨领域评估
我们提出DAHL,一个针对生物医学领域长文本生成中幻觉现象的基准数据集与自动化评估系统。该数据集从生物医学研究论文中精心构建,包含8573个问题,覆盖29个类别。DAHL通过将模型回答分解为基本信息单元,分别评估其准确性并取平均值,形成DAHL分数,相比依赖多选题的传统方法更具深度。在8种不同模型上的实验表明,更大模型幻觉更少;但当模型参数量超过70亿至80亿时,进一步扩大规模对事实准确性提升有限。DAHL分数有望替代人工标注偏好标签,且可拓展至其他专业领域。相关数据集与代码已公开。
原文摘要 · Abstract (English)
We introduce DAHL, a benchmark dataset and automated evaluation system designed to assess hallucination in long-form text generation, specifically within the biomedical domain. Our benchmark dataset, meticulously curated from biomedical research papers, consists of 8,573 questions across 29 categories. DAHL evaluates fact-conflicting hallucinations in Large Language Models (LLMs) by deconstructing responses into atomic units, each representing a single piece of information. The accuracy of these responses is averaged to produce the DAHL Score, offering a more in-depth evaluation of hallucinations compared to previous methods that rely on multiple-choice tasks. We conduct experiments with 8 different models, finding that larger models tend to hallucinate less; however, beyond a model size of 7 to 8 billion parameters, further scaling does not significantly improve factual accuracy. The DAHL Score holds potential as an efficient alternative to human-annotated preference labels, being able to be expanded to other specialized domains. We release the dataset and code in public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。