arXiv:2606.21875stat.MLcs.LG2026-06

提出可检测证据冲突与稳定性的审计工具SEF,让模型判断更透明。

Signed Evidence Flow: Conflict-Aware and Stability-Calibrated Data Analysis

论文配图:Signed Evidence Flow: Conflict-Aware and Stability-Calibrated Data Analysis
图 1 · 摘自论文原文
  • 用符号化特征归因追踪支持与反对证据,量化冲突与稳定性
  • 在多个数据集上发现高置信度预测中仍存在风险差异,冲突能提升风险排序
  • 引入ScopeGate诊断方向,适合作为模型审计而非通用风险评分

现代数据分析常给出置信度但不揭示证据是否清晰、冲突或稳定。同一置信度下,可能一者证据一致,另一者则支持与反对并存。本文提出签名证据流(SEF),结合拟合预测规则与带符号的特征归因,测量支持、反对、冲突及扰动稳定性。证明当置信度同时决定总证据质量时,其可精确衡量冲突;推导剩余条件方差,并说明冲突在某些情况下可超越置信度与其他审计变量提升损失预测能力。还建立冲突与几何决策脆弱性的关联。在医疗、Covertype、黑箱、金融等10个外部数据集上,冲突能区分看似可信预测中的真实风险。交叉验证显示,其在多个数据集(包括两个大型金融任务)中提供额外误差排序信息。方向非普适:部分任务中低冲突反而更危险。因此引入霍尔德外置换诊断器ScopeGate,在使用SEF前验证方向有效性。故SEF是审计工具,描述证据结构,而独立校准样本决定该结构在目标群体中是否有效。

原文摘要 · Abstract (English)

Modern data analysis usually gives a prediction without showing whether the evidence behind it is clear, conflicting, or stable. Two cases can have the same fitted confidence even when one has mostly agreeing evidence and the other has strong support and strong opposition. We propose Signed Evidence Flow (SEF), which combines a fitted prediction rule with signed feature attributions to measure support, opposition, conflict, and perturbation stability. We prove that confidence determines conflict exactly when it also determines total evidence mass, derive the remaining conditional variance, and state when conflict can improve loss prediction beyond confidence and other audit variables. We also connect conflict to geometric decision fragility. Across healthcare, Covertype, black-box, finance, and ten external data sets, conflict sometimes separates risk among predictions that already appear confident. Cross-fitted tests show added error-ranking information beyond confidence and attribution entropy on several data sets, including two large finance tasks. The direction is not universal: in some tasks, lowconflict cases are riskier. We therefore introduce ScopeGate, a held-out permutation diagnostic that checks the direction before SEF is used for review triage. SEF is consequently an audit tool rather than a universal risk score: it describes evidence structure, while an independent calibration sample determines whether that structure is useful in the target population.

模型审计证据冲突风险评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。