arXiv:2608.11267eess.IVcs.CV2026-08

发现表征可解码但不可靠,几何结构变化不影响可靠性评估效果

Decodable but Not Accessible: Auditing Distance-Based Reliability Estimation on Disentangled Skin-Lesion Representations

论文配图:Decodable but Not Accessible: Auditing Distance-Based Reliability Estimation on Disentangled Skin-Lesion Representations
图 1 · 摘自论文原文
  • 通过调整正交性强度改变表征几何结构,测试其对可靠性估计的影响
  • 多种距离度量方法在各水平下表现均低于随机(约0.40 AUROC)
  • 信息仍可被探测但无法被非探针方法利用,揭示可解码≠可访问

基于距离的可靠性估计假设表征几何反映可信度,但该假设极少在直接重塑几何的训练干预下被检验。本文通过域对抗表示学习中的解耦剂量-反应阶梯进行审计,三组检查点具有相同架构与16维表征,仅正交性强度(lambda = 0, 1, 5)不同。随着解耦强度提升,表征几何发生显著变化:条件数变化达两个数量级(肯德尔tau = 0.84,p = 2.8e-5)。然而,这一变化未带来可靠性估计改善:马氏距离AUROC(ISIC-test vs. PAD-UFES)在所有水平均保持平稳且低于随机水平(约0.40),且与五种几何度量均无显著关联。余弦距中心、池化k近邻等距离类评分器,以及能量置信度、虚拟对数匹配、核密度估计等非距离类评分器也表现出类似失败。八种评分器中有七种结果一致;唯一出现上升趋势的能量评分不构成整体反例。一个无需访问训练目标的监督探针在所有水平上均以0.72–0.81 AUROC恢复域归属,表明相关信息并未消失。研究揭示:分类性能优异并不意味着信息组织形式适合下游可靠性估计器使用,信息可能可解码却对非探针评估器基本不可访问。

原文摘要 · Abstract (English)

Distance-based reliability estimation assumes that a representation's geometry reflects its trustworthiness, yet this assumption is rarely tested under training interventions that reshape geometry directly. We audit this assumption under domain-adversarial representation learning using a disentanglement dose-response ladder. Three checkpoint families share the same architecture and a 16-dimensional representation, differing only in orthogonality strength (lambda = 0, 1, 5). Representation geometry changed substantially with disentanglement strength: the condition number shifted by two orders of magnitude (Kendall tau = 0.84, exact p = 2.8e-5). This change was not accompanied by improved reliability estimation: Mahalanobis-distance AUROC (ISIC-test vs. PAD-UFES) remained flat and below chance (about 0.40) at every level, with no significant association with any of five geometry metrics tested. The same failure was observed for cosine-to-centroid and pooled k-nearest-neighbor scorers, plus three non-distance-based scorers: an energy-based confidence score, Virtual-Logit Matching, and a kernel density estimator. Seven of eight scorers converged on the same result; the energy-based score showed an isolated upward trend that we report but do not treat as evidence against the overall pattern. A supervised probe with no access to the training objective recovered domain membership from the identical embeddings at 0.72-0.81 AUROC across every level, showing that the relevant information was not absent from the representation. These findings indicate that classification performance alone can overlook whether information in a learned representation is organized in a form that downstream reliability estimators can use. Information can remain decodable while becoming largely inaccessible to non-probing reliability estimators.

表示学习可靠性评估解耦表征可访问性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。