构建多语言科学文本幻觉检测数据集,助力大模型可信生成
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
- 从16个公开模型生成答案,覆盖5种高资源与4种低资源语言
- 含900个科学问题与7000+条回答,标注事实性错误与语言流畅性
- 专为科研场景设计,适合评估大模型幻觉与跨语言可靠性
我们提出CAP(Confabulations from ACL Publications)数据集,一个用于研究大语言模型在科学文本生成中幻觉现象的多语言资源。该数据集聚焦科学领域,因术语专业、统计推理复杂且语境依赖性强,易导致事实性错误。现有模型缺乏真正理解能力,常产生表面化泛化偏差。CAP涵盖五种高资源语言(英语、法语、印地语、意大利语、西班牙语)和四种低资源语言(孟加拉语、古吉拉特语、马拉雅拉姆语、泰卢固语),包含900个精心筛选的科学问题及超过7000条来自16个公开模型的生成回答,以问答对形式提供,附带标记序列与对应logits。每条实例均标注二元标签:是否含科学幻觉(即事实性错误)以及语言流畅性标签,反映文本质量或自然度问题。数据集已公开,旨在推动幻觉检测、多语言评估与更可靠的科学NLP系统研发。
原文摘要 · Abstract (English)
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text generation. CAP focuses on the scientific domain, where hallucinations can distort factual knowledge, as they frequently do. In this domain, however, the presence of specialized terminology, statistical reasoning, and context-dependent interpretations further exacerbates these distortions, particularly given LLMs' lack of true comprehension, limited contextual understanding, and bias toward surface-level generalization. CAP operates in a cross-lingual setting covering five high-resource languages (English, French, Hindi, Italian, and Spanish) and four low-resource languages (Bengali, Gujarati, Malayalam, and Telugu). The dataset comprises 900 curated scientific questions and over 7000 LLM-generated answers from 16 publicly available models, provided as question-answer pairs along with token sequences and corresponding logits. Each instance is annotated with a binary label indicating the presence of a scientific hallucination, denoted as a factuality error, and a fluency label, capturing issues in the linguistic quality or naturalness of the text. CAP is publicly released to facilitate advanced research on hallucination detection, multilingual evaluation of LLMs, and the development of more reliable scientific NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。