机器提取法律逻辑时出错,这篇论文给出验证其可信度的方法。
When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic
- 用蒙特卡洛模拟测试不同提取器的分歧,筛选出稳定成立的法律推论
- 在密苏里州29,365个条款中,93.2%的章节因误差过高不达标
- 适合法律与AI交叉研究者,尤其关注自动化法规解析可靠性的人
随着机器解析法律条文日益普遍,不同提取工具间存在分歧:在密苏里州法律中,两个独立提取器对数值阈值的存在性判断错误率高达0.43。本文探讨何种形式逻辑能在这种噪声下存活。构建了针对杜昆-盖吉蕴含基的被动生存证书:测量各属性在提取器间的差异,通过1,000次蒙特卡洛试验回放,仅当单侧威尔逊95%置信下界超过0.95时,才认证一条蕴含关系;每条被认证的推论均附带前提片段和最小反例。在29,365个密苏里州条款及502个印度中央法案条款上,预注册的保留门控测试通过(7个标题下10个法系精确匹配;11个标题下16个含5%容差),但在全局部署的误差模型下,93.2%的保留章节低于信息量底线,2×2因子分析表明问题源于校准率迁移而非选择偏差。该证书可用但脆弱,建议按章节校准或容忍误差。代码、数据产品与审计日志(含一项撤回声明)已公开。
原文摘要 · Abstract (English)
Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri's statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43. We ask what formal logic survives such noise. We build a passive survival certificate for the Duquenne-Guigues implication basis of machine-extracted statutory contexts: per-attribute inter-extractor disagreement is measured, replayed against the basis in 1,000 Monte Carlo trials, and an implication is certified only when a one-sided Wilson 95% lower bound on survival reaches 0.95; every certified implication carries premise spans and a minimal counterexample. On 29,365 Missouri sections and 502 Indian central-Act sections, the preregistered held-out gate passes (10 statute families across 7 Titles exact; 16 across 11 with 5% tolerance), yet under one globally deployed error model 93.2% of held-out chapters fall below the informativeness floor, and a 2x2 factorial assigns that to calibration-rate transfer, not selection. The certificate is usable but fragile: deploy it per-chapter-calibrated or error-tolerant. Code, data products, and the audit trail, including one retracted claim, are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。