检测法国法院判决中未明示的法律条文引用,发现专家分歧处模型最易出错。
Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions
- 通过专家标注的1015对文本构建基准,识别隐含法律引用。
- 模型在专家有分歧的案例上错误率高,占所有误报的三分之二。
- 用多模型共识排序可实现无监督下前200候选76%精确率。
大规模应用计算方法于法律领域需区分真正的法律推理与表面相似性。我们以检测法国《民法典》的隐含引用为例,即法院适用法律条文但未明确提及。为此,我们发布了一个由三位法律专家标注的1,015个段落-文章配对的基准数据集。核心发现是:专家分歧本身具有信息量——三分之一存在争议的案例正是模型失败之处。最佳集成模型整体F1得分为0.70,但其三分之二的假阳性出现在这些争议案例中,该现象在评估的十种模型中均成立。专家分歧反映的是内在难度,而非标注噪声。这并不妨碍实用工具的开发:若将任务重构为多模型共识的前k名排序,无需监督即可在前200个候选中达到76%的精确率。
原文摘要 · Abstract (English)
Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a concrete task: detecting implicit citations of the French Civil Code, where a court applies a statutory rule without naming it: a post-hoc question about the reasoning a court actually used. We release a benchmark of 1,015 passage-article pairs annotated by three legal experts. Our central finding is that their disagreement is itself informative: the third of cases the experts dispute are where models fail. Our best ensemble reaches an F1 score of 0.70 overall. Yet, two-thirds of its false positives fall on those disputed cases, a concentration that holds across all ten models we evaluate. Disagreement is a signal of intrinsic difficulty, not annotation noise. This should not block useful tools, however: reframed as top-$k$ ranking with multi-model consensus, the same signals reach 76% precision for the top-200 candidates without supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。