法律检索基准中,罪名匹配可解释大部分性能差异,需警惕其对评估结果的干扰
Charge as a Construct-Validity Factor in Chinese Legal Case Retrieval: A Cross-Benchmark Audit
- 用罪名统一排序+BM25微调,可复现99.2%的模型性能差距
- 罪名相同则案例相关性极高(宏AUC 0.871),但模型真正优势仅剩小部分
- 提出可复用的罪名控制协议,帮助未来评估避免此类偏差
中文法律案例检索基准根据案件定性是否与查询匹配来评分,当前最优系统在LeCaRDv2上达到NDCG@10为0.85-0.88。多数从BM25到最佳训练模型的差距,可通过仅按共享主要罪名排序、再以BM25细分解决,覆盖99.2%的差距,且与最优系统无显著差异。这反映基准设计本质:最高相关性由罪名的核心构成要素定义,因此同罪名案例天然相关(相关性提升4.49;罪名到相关性宏AUC 0.871)。固定罪名后,训练重排序器相较于BM25的优势几乎消失(+0.026 NDCG@10,置信区间不含零,约四分之一)。该效应不一致:在LeCaRDv1上恢复84.3%,但在CAIL2022上失效,罪名相关信号逐步减弱(宏AUC 0.871/0.759/0.728);预测罪名级联可复现LeCaRDv2上76.6%的性能,但无法迁移。罪名影响也体现在初筛阶段:零训练的罪名池通道提升LeCaRDv2召回率(R@100 +0.025),错误罪名对照组表现更差,作为混淆因素的正向控制,非新方法或创新主张。因此,罪名是基准层面的关键构念有效性因子——非统一解释指标,也不代表系统依赖罪名。我们构建了可复用的罪名控制协议(CCE),在三个基准上触发均为零或描述性结果,符合预期。代码、模式和协议已发布,供未来基准筛选,防止将NDCG@10误读为法律推理能力。
原文摘要 · Abstract (English)
Chinese Legal Case Retrieval (LCR) benchmarks grade a reference judgment relevant when its legal characterization matches the query, and strong systems now reach NDCG@10 of 0.85-0.88. Most of the BM25-to-best-trained gap is recoverable with no retrieval model: ranking candidates only by shared primary charge, broken by BM25, closes 99.2% of it on LeCaRDv2 -- with no detectable difference from the best-trained system. This reflects benchmark design: LeCaRDv2 defines top relevance via the crime's key constitutive elements, which encode the charge, so same-charge cases are relevant by construction (relevance lift 4.49; charge-to-relevance macro-AUC 0.871). Holding charge fixed, the trained reranker's advantage over BM25 collapses to a small within-charge residual (+0.026 NDCG@10, cluster-bootstrap CI excluding zero, about a quarter), the only non-definitional positive. The effect is not uniform: the same rule recovers 84.3% on LeCaRDv1 and is out of spec on CAIL2022, with the charge-to-relevance signal weakening in step (macro-AUC 0.871/0.759/0.728); a predicted-charge cascade reproduces 76.6% on LeCaRDv2 but does not transfer. The construct is also cashable at first stage: an exploratory zero-training charge-pool channel lifts LeCaRDv2 recall (R@100 +0.025, wrong-charge controls hurt), reported as a positive control for the confound, not a retrieval method or novelty claim. Charge is thus a high-leverage construct-validity factor at the benchmark level -- not auniform explanation of NDCG@10, and not evidence that any system relies on charge. We package established construct-validity and partial-input checks as a reusable charge-controlled protocol (CCE); on all three benchmarks its triggers come back null or descriptive, behaving as designed. We release the scripts, schema, and protocol so future benchmarks can be screened before their NDCG@10 is read as legal-reasoning ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。