提出文档抽取的可信风险控制框架,解决三类失效问题。
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
- 构建有效性阶梯,分层修复风险控制失效问题
- 在真实数据上实现0.096风险下0.318覆盖率,接近理论最优
- 适合追求高可靠性文档系统开发与验证的研究者
针对文档抽取中按字段选择性控制风险的需求,本文诊断出三大失效模式:文档聚类(设计效应1.84-2.45)、分数重拟合泄漏(在风险0.127时覆盖仅0.416,违反α=0.10的情况达95%分割)以及阈值网格坍塌的纠缠病态(从0.030降至0.001)。通过构建有效性阶梯,逐级保障风险控制。采用训练/验证分割协议恢复学习融合模型的预期选择性风险控制,实现风险0.096、覆盖0.318,名义α=0.10;但实际风险超限比例达47.5%,非严格证书。使用曼德里安学习后测试结合精确二项尾部,获得组内独立的PAC证书:场间独立0.171风险0.068,聚类校正0.140,文档独立0.060——唯一与真实文档匹配的可信层级。支持箱(pre-specified provenance taxonomy)在该数据集表现显著优于其他方法(p<1e-4,Bonferroni校正),但在haiku或qwen上未复现;高精度语料中池化阈值更优,条件化在无法认证时仍有效。冻结配置验证在未触碰选择的claude-haiku-4-5上于双风险水平均成立,且三标注员人工黄金审计确认接受集风险为1.3%,低于10%预算(Fleiss' kappa=0.83;标签单向保守)。代码已开源,采用种子锁定与回归门控流程。
原文摘要 · Abstract (English)
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real documents. On 13,859 genuine claude-sonnet-5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes: document clustering (design effect 1.84-2.45), score-refit leakage (coverage 0.416 at risk 0.127, violating alpha=0.10 in 95% of splits), and a tie-mass pathology (a degenerate score collapses the threshold grid, 0.030 to 0.001). We organize the fixes as a validity ladder, guarantee form stated per tier. A fit/val split protocol restores expected-selective-risk control for a learned fusion: coverage 0.318 at risk 0.096 at nominal alpha=0.10, no tolerance band (production variant 0.326) -- an on-average point whose realized risk exceeds alpha in 47.5% of resplits, not a certificate. Mondrian Learn-then-Test with exact binomial tails yields per-group PAC certificates: field-iid 0.171 at risk 0.068, cluster-corrected 0.140, doc-iid 0.060 -- the only tier matching documents, honestly near-vacuous today. Support-bin, the pre-specified provenance taxonomy, wins every rigor tier on the sonnet CORD capture (p<1e-4, Bonferroni-corrected) -- a win that does not replicate on the same documents under haiku or qwen -- while on higher-accuracy corpora pooled thresholds win: conditioning helps exactly where pooled cannot certify, subsumed by a learned score elsewhere. A frozen-configuration confirmation on selection-untouched claude-haiku-4-5 held at both risk levels, and a blind three-annotator human-gold audit verifies the practical tier's accepted-set risk at 1.3% against its 10% budget (Fleiss' kappa=0.83; labels err one-sidedly pessimistic). Released Apache-2.0 with seed-pinned, regression-gated procedures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。