提出可复现的专家参与型幻觉评测框架,提升大模型评估可信度。
The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
- 构建可重复、开放且结合领域背景的幻觉评测体系
- 实证显示缺乏专家参与会导致评测结果无效
- 适合模型评估者与安全研究人员参考
模型生成文本中看似合理却错误的内容被普遍认为广泛存在,对大模型的负责任应用构成挑战。然而,目前尚缺乏系统性的科学工作来全面衡量语言模型幻觉的普遍性。本文主张应采用可复现、开放且基于具体领域语境的幻觉评测方法。我们提出一个幻觉分类体系,并通过案例研究证明:若在数据生成初期缺少领域专家参与,所获幻觉指标将缺乏有效性与实际应用价值。
原文摘要 · Abstract (English)
Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern, there is little scientific work that attempts to measure the prevalence of language model hallucination in a comprehensive way. In this paper, we argue that language models should be evaluated using repeatable, open, and domain-contextualized hallucination benchmarking. We present a taxonomy of hallucinations alongside a case study that demonstrates that when experts are absent from the early stages of data creation, the resulting hallucination metrics lack validity and practical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。