用人类编写的幻觉样本替代模型生成,让评估更持久可靠。
Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

- 用人工撰写1600个幻觉样本,覆盖四种语言
- 人工样本与模型样本分布相似,检测效果相当
- 适合需要长期稳定评估幻觉的科研与产品团队
在模型快速迭代的时代,如何让幻觉评估更具持久性?我们探索用人类编写的幻觉样本替代模型生成的样本,使检测评估摆脱对特定模型的依赖。为此,我们构建了一个包含1,600个由人类撰写的样本的数据集,覆盖中文、英文、法文、意大利文四种语言,并收集了来自五种视觉-语言模型的18,400个样本,均采用细粒度的跨度级标注方式标注幻觉。结果表明,人类编写的样本具有更高标注一致性,能更好控制数据内容,同时在分布上仍与模型生成样本保持相似,且能合理反映检测能力,说明人类数据可作为模型幻觉基准的有效替代方案。
原文摘要 · Abstract (English)
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。