提出新评估框架CiteEval,更精准衡量论文引用质量。
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
- 基于引文原则构建细粒度评估体系,涵盖查询、检索上下文等
- 建立多领域人工标注数据集CiteBench,支持高质量引文评估
- 推出自动化评估工具CiteEval-Auto,与人工判断高度一致
引文质量直接影响信息获取系统的可信度与效果。现有评估框架多依赖自然语言推理(NLI)进行二元或三元支持性判断,我们认为这并非最优代理指标。本文提出基于原则的引文评估框架CiteEval,实现对引文在广泛上下文(包括引用来源、完整检索上下文、用户查询和生成文本)中的细粒度评估。基于该框架,构建了多领域基准CiteBench,包含高质量人工标注的引文质量标签。为实现高效评估,进一步开发了与人类判断高度相关的模型化评估指标CiteEval-Auto。在多种系统上的实验表明,相比现有方法,CiteEval-Auto能更好捕捉引文的多维度特性,提供一种原理清晰且可扩展的模型生成引文评估方案。
原文摘要 · Abstract (English)
Citation quality is crucial in information-seeking systems, directly influencing trust and the effectiveness of information access. Current evaluation frameworks, both human and automatic, mainly rely on Natural Language Inference (NLI) to assess binary or ternary supportiveness from cited sources, which we argue is a suboptimal proxy for citation evaluation. In this work we introduce CiteEval, a citation evaluation framework driven by principles focusing on fine-grained citation assessment within a broad context, encompassing not only the cited sources but the full retrieval context, user query, and generated text. Guided by the proposed framework, we construct CiteBench, a multi-domain benchmark with high-quality human annotations on citation quality. To enable efficient evaluation, we further develop CiteEval-Auto, a suite of model-based metrics that exhibit strong correlation with human judgments. Experiments across diverse systems demonstrate CiteEval-Auto's superior ability to capture the multifaceted nature of citations compared to existing metrics, offering a principled and scalable approach to evaluate and improve model-generated citations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。