arXiv:2502.10881cs.CL2025-02被引 2

首个中文引文忠实性检测数据集,解决标注成本高难题

CiteCheck: Towards Accurate Citation Faithfulness Detection

  • 两阶段人工标注+大模型生成负样本,降低标注成本
  • 测试集难度高,顶尖大模型准确率仍不足
  • 小模型经高效微调即可达到强性能,适合中文RAG研究

引文忠实性检测对提升检索增强生成(RAG)系统至关重要,但中文大规模数据集稀缺。现有方法因需人工标注负样本而成本高昂。为此,我们提出首个大规模中文数据集CiteCheck,采用两阶段人工标注的低成本构建方式,在平衡正负样本的同时显著降低标注开销。CiteCheck包含训练与测试集。实验表明:(1)测试样本极具挑战性,即使最先进的大模型也难以达到高准确率;(2)使用大模型生成的负样本扩充训练数据,结合参数高效微调,使小型模型也能取得优异表现。该数据集为中文RAG系统的引文忠实性检测提供坚实基础,现已公开,便于学术研究。

原文摘要 · Abstract (English)

Citation faithfulness detection is critical for enhancing retrieval-augmented generation (RAG) systems, yet large-scale Chinese datasets for this task are scarce. Existing methods face prohibitive costs due to the need for manually annotated negative samples. To address this, we introduce the first large-scale Chinese dataset CiteCheck for citation faithfulness detection, constructed via a cost-effective approach using two-stage manual annotation. This method balances positive and negative samples while significantly reducing annotation expenses. CiteCheck comprises training and test splits. Experiments demonstrate that: (1) the test samples are highly challenging, with even state-of-the-art LLMs failing to achieve high accuracy; and (2) training data augmented with LLM-generated negative samples enables smaller models to attain strong performance using parameter-efficient fine-tuning. CiteCheck provides a robust foundation for advancing citation faithfulness detection in Chinese RAG systems. The dataset is publicly available to facilitate research.

引文检测中文RAG数据集高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。