构建中文事实核查的原子化主张分解数据集,提升长文本回答可信度。
A Claim Decomposition Benchmark for Long-form Answer Verification
- 基于WebCPM构建专家标注的中文原子主张数据集
- 包含500组问答对,共4956个可验证的原子主张
- 为大模型事实核查提供新基准,适合可信AI研究者使用
大语言模型在复杂长文本问答任务中表现显著提升,但生成内容常存在幻觉问题。现有研究多关注回答的准确引用,却忽视了对回答中主张的识别。为此,本文提出新的主张分解基准,要求系统能识别大模型回答中的原子且可验证主张。我们构建了中文原子主张分解数据集(CACDD),基于WebCPM数据集并经专家标注,确保高质量。CACDD包含500组人工标注的问答对,共4956个原子主张。同时,我们设计了新的标注流程并分析任务挑战。此外,提供了零样本、少样本及微调大模型的实验结果。结果显示,主张分解极具挑战性,需进一步探索。所有代码与数据公开于https://github.com/FBzzh/CACDD。
原文摘要 · Abstract (English)
The advancement of LLMs has significantly boosted the performance of complex long-form question answering tasks. However, one prominent issue of LLMs is the generated "hallucination" responses that are not factual. Consequently, attribution for each claim in responses becomes a common solution to improve the factuality and verifiability. Existing researches mainly focus on how to provide accurate citations for the response, which largely overlook the importance of identifying the claims or statements for each response. To bridge this gap, we introduce a new claim decomposition benchmark, which requires building system that can identify atomic and checkworthy claims for LLM responses. Specifically, we present the Chinese Atomic Claim Decomposition Dataset (CACDD), which builds on the WebCPM dataset with additional expert annotations to ensure high data quality. The CACDD encompasses a collection of 500 human-annotated question-answer pairs, including a total of 4956 atomic claims. We further propose a new pipeline for human annotation and describe the challenges of this task. In addition, we provide experiment results on zero-shot, few-shot and fine-tuned LLMs as baselines. The results show that the claim decomposition is highly challenging and requires further explorations. All code and data are publicly available at \url{https://github.com/FBzzh/CACDD}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。