arXiv:2607.18278cs.LGcs.AI2026-07

发现并定位模型错误但高自信的集中区域,提升校准可靠性。

FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration

论文配图:FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration
图 1 · 摘自论文原文
  • 通过置信度、局部支持等多信号融合,后验排序预测风险。
  • 在7个数据集上,该方法比基准高出30%以上,能捕获大部分危险错误。
  • 适用于需识别高危预测区的可信系统,如医疗或金融决策。

校准通常以整体评估为主,但最危险的失败往往是局部性的:模型在错误时仍保持高度自信。本文将这种现象称为‘虚假自信集中’,即高自信错误在预测空间中占据紧凑且可发现的区域。提出FALCON-Discover,一种后验、模型无关的框架,利用置信度、局部支持、邻域一致性与扰动稳定性等差异信号对预测进行排序。在7个二分类表格数据集、4个种子、五折交叉拟合及强学习器(如XGBoost和CatBoost)下,发现虚假自信集中具有重复性但依赖于具体场景。在主置信度阈值下,基于差异的排序显著优于最优验证选择的校准或信任评分基线,在最强场景下表现更优;而原始置信度仅捕获极少危险错误质量。最佳检测方式因数据集而异:当需融合多个线索时,学习型差异表现最强;当局部决策脆弱性主导时,稳定性中心排序效果最佳。结果表明,危险过自信应视为家族级发现任务,而非单一分数校准问题,并推动针对置信度、支持与稳定性分歧区域的校准策略。

原文摘要 · Abstract (English)

Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong. We study this failure mode as false-confidence concentration, the extent to which confident errors occupy compact, discoverable regions of prediction space. We introduce FALCON-Discover, a post-hoc, model-agnostic framework that ranks predictions using discrepancy signals from confidence, local support, neighborhood agreement, and perturbation stability. Across seven binary tabular datasets, four seeds, five-fold cross-fitting, and strong learners including XGBoost and CatBoost, we find that false-confidence concentration is recurrent but regime-dependent. At the main confidence threshold, discrepancy-based ranking substantially outperforms the strongest validation-selected calibration or trust-scoring baseline in the strongest regimes, while raw confidence recovers little dangerous-error mass. The best detector varies across datasets: learned discrepancy is strongest when multiple cues must be combined, whereas stability-centered ranking works best when local decisional fragility dominates. These results show that dangerous overconfidence is better treated as a family-level discovery problem than as a single-score calibration problem, and motivate calibration strategies that explicitly target regions where confidence, support, and stability diverge.

模型校准异常检测置信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。