用分布内数据预判干扰检测器在真实场景中的性能下降
Anticipating the Optimism Gap: Predicting Distribution-Shift Degradation of RF-Impairment Detectors from In-Distribution Statistics
- 基于分布内统计量构建岭回归模型预测性能衰减
- 实测数据中预测效果显著,最高相关性达1.0
- 适用于评估卫星导航抗干扰检测器的可靠性
GNSS射频干扰(欺骗、干扰、多径)检测器通常仅报告其训练分布上的单个AUC值。当环境变化时,该指标会下降,但降幅难以提前预知,因带标签的真实数据稀缺。本文提出能否在未见域外数据前就预测这种性能乐观偏差。在可调严重度的参数化合成测试平台上,评估了13种检测器(5种物理基线、全特征逻辑回归与多层感知机、单特征学习控制),发现性能差距随偏离程度单调增长(均斯皮尔曼相关0.50)。该差距主要由检测器使用可观测变量数量决定,而非是否为学习模型,且不同干扰类型差异明显。关键发现:仅用分布内得分统计构建的岭回归模型,能有效预测未见过的检测器(R²=0.47)和未见过的干扰类型(R²=0.46),且经2000倍置换检验显著(p<0.001),即使移除目标特征仍稳健。合成结果后,在三个公开野外数据集上验证:在Jammertest 2024中跨检测器预测成立(R²=0.11, p=0.009);在SatGrid中,高严重度下真实AUC比分布内值低最多0.22,甚至导致符号反转,且两者秩相关系数达1.0。该机制在真实数据中依然存在,但幅度较小。研究开源测试平台、软件接收前端、数据接入适配器及实验协议。
原文摘要 · Abstract (English)
Detectors for GNSS radio-frequency impairments (jamming, spoofing, multipath) are usually reported with a single AUC measured on the distribution they were tuned on. That number falls once conditions move, and the size of the drop is rarely known in advance because labelled field data is scarce. We ask whether this optimism can be predicted before any out-of-distribution data is seen. On an open, parameter-grounded synthetic testbed with a tunable severity shift, we evaluate thirteen detectors (five physics baselines, full-feature logistic regression and multilayer perceptrons, and single-feature learned controls) across four impairment classes. The optimism gap, the difference between in-distribution and shifted AUC, grows monotonically as the shift deepens (mean Spearman correlation 0.50). It is driven by how many observables a detector uses rather than by whether it is learned, and it varies systematically by class. Centrally, a ridge model built only from in-distribution score statistics predicts the gap for a detector it has never seen (R^2 = 0.47) and for an impairment class it has never seen (R^2 = 0.46); both are significant against a 2000-fold permutation null (p < 0.001) and survive removing the feature that is, by construction, part of the target. The headline findings are synthetic. We then run the pre-registered protocol on three open field corpora: on Jammertest 2024 the cross-detector prediction holds (R^2 = 0.11, p = 0.009), and on SatGrid, whose spoofer power sweep gives a calibrated severity axis, in-distribution AUC overstates higher-severity AUC by up to 0.22 and to the point of sign inversion, with in-distribution AUC and realised gap perfectly rank-correlated (Spearman rho = 1.0). The mechanism survives contact with real data, at smaller magnitude than in simulation. We release the testbed, a software-receiver front end, the ingest adapters and the protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。