提出评估合成数据在长文本事实核查中作用的框架,提升模型可解释性。
SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- 构建三维度评估框架:输入特性、生成逻辑与解释质量
- 合成数据能提升基础模型验证性能,尤其结合人工数据时效果更优
- 即使验证准确率未提升,也能改善模型解释一致性,适合可信AI研究者
大语言模型(LLMs)具备长上下文处理能力,可直接推理长文档,减少分块或检索需求。然而,构建标注数据集成本高昂。合成数据提供了可扩展的替代方案。本文提出 SynClaimEval 框架,用于评估合成数据在长上下文事实核查任务中的有效性——该任务对幻觉检测和事实核验至关重要。框架从三个维度展开:(i) 输入特征,通过改变上下文长度并测试跨域泛化能力;(ii) 生成逻辑,控制论断复杂度与错误类型变化;(iii) 解释质量,衡量模型解释是否与预测结果一致。跨多个基准测试的实验表明,长上下文合成数据可提升指令微调模型的验证表现,尤其在增强现有人工数据集时效果显著。此外,合成数据还能提升解释质量,即使验证得分未提高,也凸显其在提升性能与可解释性方面的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) with extended context windows promise direct reasoning over long documents, reducing the need for chunking or retrieval. Constructing annotated resources for training and evaluation, however, remains costly. Synthetic data offers a scalable alternative, and we introduce SynClaimEval, a framework for evaluating synthetic data utility in long-context claim verification -- a task central to hallucination detection and fact-checking. Our framework examines three dimensions: (i) input characteristics, by varying context length and testing generalization to out-of-domain benchmarks; (ii) synthesis logic, by controlling claim complexity and error type variation; and (iii) explanation quality, measuring the degree to which model explanations provide evidence consistent with predictions. Experiments across benchmarks show that long-context synthesis can improve verification in base instruction-tuned models, particularly when augmenting existing human-written datasets. Moreover, synthesis enhances explanation quality, even when verification scores do not improve, underscoring its potential to strengthen both performance and explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。