arXiv:2602.24278cs.LG2026-02被引 1

现有表征可识别性评估方法在特定条件下才可靠,否则会误判。

Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations

  • 区分数据生成过程与编码器结构假设,揭示评估失效根源
  • 发现经典指标在真实场景中常产生误报和漏报
  • 提供可复现的测试工具,适合研究表征学习可信度者使用

表征学习中的可识别性通常通过标准指标(如MCC、DCI、R²)在已知真实因子的合成基准上评估。这些指标被假定能反映理论保证的等价类内恢复能力。我们发现该假设仅在特定结构条件下成立:每种指标都隐含地依赖于数据生成过程(DGP)和编码器的假设。当这些假设不成立时,指标会失准,导致系统性的假阳性与假阴性。此类失败既出现在经典可识别性范式中,也发生在最需要可识别性验证的后验场景中。本文提出一个分类框架,分离DGP假设与编码器几何特性,据此刻画现有指标的有效范围,并发布一个评估套件用于可复现的压力测试与对比。

原文摘要 · Abstract (English)

Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with known ground-truth factors. These metrics are assumed to reflect recovery up to the equivalence class guaranteed by identifiability theory. We show that this assumption holds only under specific structural conditions: each metric implicitly encodes assumptions about both the data-generating process (DGP) and the encoder. When these assumptions are violated, metrics become misspecified and can produce systematic false positives and false negatives. Such failures occur both within classical identifiability regimes and in post-hoc settings where identifiability is most needed. We introduce a taxonomy separating DGP assumptions from encoder geometry, use it to characterise the validity domains of existing metrics, and release an evaluation suite for reproducible stress testing and comparison.

表征学习可识别性评估方法模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。