论文发现:过细的引用反而降低生成质量,中等粒度最有效。
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
- 对比句、段、文档三级引用粒度,发现段落级最优
- 细粒度引用使溯源准确率下降16%-276%
- 大模型更依赖多句信息融合,细粒度会破坏其能力
引用粒度——是引用单个句子、段落还是整篇文档——是带引用生成中的关键设计选择。尽管细粒度引用常被认为更利于人工验证,但其对模型性能的影响仍不明确。我们分析了四个规模(8B-120B)的模型,发现强制细粒度引用会使溯源质量下降16%-276%,相较最优粒度表现更差。观察到溯源质量在中等粒度(段落级)达到峰值。分析表明,细粒度引用破坏了将证据与回答主张关联所需的语义依赖关系,而过于粗粒度(多段)则引入干扰噪声。重要的是,这一性能差距随模型规模非单调变化:细粒度约束对大模型的惩罚尤为严重,说明原子级引用破坏了大模型擅长的跨句信息整合能力。令人惊讶的是,采用最优引用粒度可显著提升溯源质量,同时保持或改善答案正确性。总体而言,仅以人工验证为目标追求细粒度引用,忽视了模型自身约束,损害了溯源真实性和生成可靠性。有效的溯源需将引用粒度与模型自然语义范围对齐。
原文摘要 · Abstract (English)
Citation granularity - whether to cite individual sentences, paragraphs, or documents - is a critical design choice in attributed generation. While fine-grained citations are often preferred for precise human verification, their impact on model performance remains under-explored. We analyze four model scales (8B-120B) and demonstrate that enforcing fine-grained citations degrades attribution quality by 16-276% compared to the best-performing granularity. We observe a consistent performance pattern where attribution quality peaks at intermediate granularities (paragraph-level). Our analysis suggests that fine-grained (sentence-level) citations disrupt necessary semantic dependencies for attributing evidence to answer claims, while excessively coarse citations (multi-paragraph) introduce distracting noise. Importantly, the magnitude of this performance gap varies non-monotonically with model scale: fine-grained constraints disproportionately penalize larger models, suggesting that atomic citation units disrupt the multi-sentence information synthesis at which these models excel. Strikingly, citation-optimal granularity leads to substantial gains in attribution quality while preserving or even improving answer correctness. Overall, our findings demonstrate that optimizing solely for human verification via fine-grained citation disregards model constraints, compromising both attribution faithfulness and generation reliability. Instead, effective attribution requires aligning citation granularity with the model's natural semantic scope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。