冻结编码器异常检测的可复现性审计,发现所谓跨域迁移实为初始化效应。
Auditing Frozen-Encoder Anomaly Detection Across Mechanical Systems: Representation Provenance, Calibration, and Protocol Effects
- 通过检查模型权重加载方式,发现嵌套状态与外层字典加载差异显著。
- 近零向量表现与原始结果高度一致(相关性0.987,AUC接近0.98)。
- 性能看似跨域迁移,实为架构与评分机制导致的假象,适合关注可复现性的研究者。
本版本报告了对第一版中冻结编码器实验的可复现性审计。数值判别结果可从保存的产物中复现,但其原归因于干涉测量预训练的说法不成立。释放的检查点包含嵌套模型状态,可完整加载;而外层检查点字典加载时,EfficientNet-B0特征堆栈几乎全部未初始化。标注为干涉测量的保留嵌入范数约为10^{-12},与全新初始化的EfficientNet-B0网络一致,且比保留的ImageNet嵌入相差超过十二个数量级。另一组独立保留的近零嵌入集产生几乎相同的IMS第四测试异常得分(r=0.987)和记录级判别能力(AUC 0.9812 vs 0.9818)。因此,撤回关于IMS性能体现引力波仪器形态先验的因果主张。我们重新分析了在匹配误报率下的控制性IMS划分,并引入多变量经典信号基线。近零表示仍保持强尾部分离,尤其在第2、4次IMS运行中,但这被解释为架构与初始化的探索性效应,耦合于马氏距离评分。另一项PRONOSTIA审计表明,原始大预警时间由寿命比例基线引起;固定时间评估下,十特征经典基线优于保留编码器得分。这些结果说明,检查点来源、有限样本校准、架构设计及目标域基线可能制造跨域迁移的假象。它们也定义了在赋予冻结表示异常分数物理意义前所需的基本控制条件。
原文摘要 · Abstract (English)
This version reports a reproducibility audit of the frozen-encoder experiments presented in version 1. The numerical discrimination results are reproducible from the preserved artifacts, but their original attribution to interferometric pretraining is not supported. The released checkpoint contains a nested model state that loads without missing parameters, whereas loading the outer checkpoint dictionary leaves almost the entire EfficientNet-B0 feature stack uninitialized. Preserved embeddings labelled as interferometric have norms of order $10^{-12}$, matching freshly initialized EfficientNet-B0 networks and differing by more than twelve orders of magnitude from the preserved ImageNet embeddings. A second, separately preserved near-zero embedding set produces almost the same IMS 4th-test anomaly scores ($r=0.987$) and record-level discrimination (AUC $0.9812$ versus $0.9818$). We therefore withdraw the causal claim that IMS performance demonstrates a morphological prior transferred from gravitational-wave instrumentation. We reanalyse the controlled IMS splits at matched observed false-positive rates and add multivariate classical signal baselines. The near-zero representations retain strong tail separation, particularly in the 2nd and 4th IMS runs, but this is now interpreted as an exploratory architecture-and-initialization effect coupled to Mahalanobis scoring. A separate PRONOSTIA audit shows that the original large warning times were induced by a lifetime-fraction baseline; under fixed-time evaluation, a ten-feature classical baseline outperforms the preserved encoder scores. These results illustrate how checkpoint provenance, finite-sample calibration, architecture, and target-domain baselines can create an appearance of cross-domain transfer. They also define the controls required before assigning physical meaning to frozen-representation anomaly scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。