arXiv:2608.01792cs.AIcs.CL2026-08

首个专为文档抽取设计的置信度校准评测基准,助你判断模型自信是否靠谱。

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction

论文配图:Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
图 1 · 摘自论文原文
  • 通过20种可控降级生成1346个文档变体,覆盖全精度范围评估置信度。
  • 发现文本+图像模态置信度更准,且模型能力比参数量更能决定置信度质量。
  • 提出可量化人工审核成本的ECARB指标,适合部署可信智能文档系统者参考。

基于视觉语言模型(VLMs)的智能文档处理(IDP)依赖可靠的置信度评分,以决定信息抽取是自动执行还是交由人工审核。现有文档评测集多集中于高质量样本,导致低准确率区域数据稀疏,难以评估置信度校准效果。本文提出首个专注于置信度校准的关键词信息抽取(KIE)评测基准ConfBench,通过对多样化文档集应用20种可控降级流程,生成1,346个变体及70,000+实体级评估,覆盖全精度范围。评测了四款专有模型与三款开源模型,在三种输入模态下使用话语化和对数概率置信度估计方法,结果表明:(i) OCR+图像模态产生更准确的置信度;(ii) 模型能力是主导因素:在Claude系列中置信度质量随能力单调提升,而跨系列时参数量预测力差;(iii) 置信度校准质量差异显著,从近乎完美到严重过自信,但每模型后处理校正仅重标置信值,不影响基于排名的操作指标;(iv) 对数概率结合首词聚合优于均值与间距聚合。此外引入ECARB指标,将区分度提升转化为人工审核预算节省。我们公开发布ConfBench,支持可信IDP应用中置信度估计器与校准方法的系统研究。

原文摘要 · Abstract (English)

Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.

文档抽取置信度校准视觉语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。