arXiv:2606.03728cs.CLcs.IR2026-06

用归因分数重排法律问答检索结果,提升引用准确性。

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

论文配图:Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA
图 1 · 摘自论文原文
  • 基于归因分数构建轻量级交叉编码器重排候选段落
  • 在AQuAECHR上显著提升引用忠实度与专家答案匹配度
  • 跨模型训练可提取通用相关性信号,适合法律AI系统优化

面向法律问答的检索增强生成系统通常基于语义相似度检索段落并输入语言模型生成引用答案。以往研究假设高分段落更可能被模型有效引用。然而,在AQuAECHR基准测试中,语义相似度与段落归因分数无相关性。基于相似度的排序在候选池中表现甚至劣于随机选择,无法有效识别黄金引用段落。为此,本文训练一个轻量级交叉编码器,以连续扰动归因分数作为信号对段落进行重排。该方法在两个语言模型和五折交叉验证下评估,显著提升了引用忠实度与专家答案对齐程度。值得注意的是,两个独立模型训练的重排器虽原始归因一致性较低,但其输出趋于一致,表明交叉编码器能降低模型特异性噪声,生成共享的相关性信号,尽管同模型重排仍更优。结果证明,扰动归因可作为实用、模型无关的引用感知检索训练信号。

原文摘要 · Abstract (English)

Retrieval-augmented generation systems for legal question answering typically retrieve passages based on semantic similarity and provide them to a language model, which then generates cited answers. Prior work assumes that highly ranked passages are most likely to be usefully cited by the model. Perturbation-based attribution methods, such as C-LIME, have been used exclusively for post-hoc explanation. However, on the AQuAECHR benchmark, semantic similarity does not correlate with passage attribution. Within a retriever's candidate pool, similarity-based ranking performs worse than random selection at surfacing gold citation paragraphs. To address this limitation, a lightweight cross-encoder is trained on continuous perturbation-based attribution scores to re-rank passages prior to generation. This approach is evaluated on the AQuAECHR benchmark, using two language models and five-fold cross-validation. The re-ranker substantially improves citation faithfulness and alignment with gold expert answers. Notably, two re-rankers trained independently on different models converge beyond their raw attribution agreement. This finding indicates that the cross-encoder reduces model-specific noise and produces a shared relevance signal that partially transfers across models, although same-model re-ranking remains more effective. These results demonstrate that perturbation-based attribution provides a practical, model-agnostic training signal for citation-aware retrieval.

法律AI检索增强归因分析引用质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。