arXiv:2607.01025physics.app-phcs.LG2026-07

RF无人机识别高准确率可能虚高,因数据泄漏导致模型‘作弊’。

How Much Do RF Drone Benchmarks Overstate? A Controlled Study and Theory of Data Leakage in UAV Signal Identification

论文配图:How Much Do RF Drone Benchmarks Overstate? A Controlled Study and Theory of Data Leakage in UAV Signal Identification
图 1 · 摘自论文原文
  • 用理论和实验揭示短段切分导致训练测试重叠,引发数据泄漏。
  • 真实场景下识别准确率从0.74暴跌至0.46,接近随机水平。
  • 适合关注无人机检测可信度的研究者与防御系统开发者。

射频(RF)感知是反无人机防御的核心手段,依赖于无人机与操控者之间的控制、遥测和视频链路。现有基于RF的无人机检测与识别报告准确率普遍很高,但多数采用将少量连续记录切分为短段进行交叉验证的方法,导致同一记录的近似片段同时出现在训练集与测试集中,产生数据泄漏。本文通过理论分析与实证研究揭示该问题。利用Cover函数计数定理,证明当独立记录数R相对于特征维度d较小时,分类器可完全记忆记录-标签映射,即当2R ≤ d时,模型天真准确率趋近1,性能膨胀量接近1 - ACC*(ACC*为贝叶斯最优准确率)。一旦R超过此可分阈值,膨胀才缓解。受控合成实验以10组种子验证预测曲线:天真平衡准确率随记录内干扰增长从贝叶斯水平升至1.0,而诚实按记录组划分的评估则降至随机水平,差距达约0.5。在公开的DroneRF数据集上,合并留一记录外交叉验证显示,AR与Bebop机型识别的宏观F1从0.74骤降至0.46(两分类随机水平)。泄漏路径消融实验表明,几乎全部性能膨胀源于段级泄漏。

原文摘要 · Abstract (English)

Radio-frequency (RF) sensing is a central modality for counter-unmanned-aerial-system (counter-UAS) defence because it exploits the control, telemetry, and video links between a drone and its operator. Reported accuracies for RF-based drone detection and identification are often very high, but many are obtained using cross-validation that splits a small number of continuous recordings into short segments. This can place near-duplicate slices of the same recording in both training and test partitions, creating data leakage. We study this leakage problem through theory and measurement. We formalise the optimism of segment-level cross-validation and show, using Cover's function-counting theorem, that a classifier can exactly memorise the recording-to-label map when the number of independent recordings, R, is small relative to the feature dimension, d. In particular, this can occur when 2R is less than or approximately equal to d. Under these conditions, naive accuracy approaches 1, and the inflation gap approaches 1 - ACC*, where ACC* is the Bayes accuracy. The inflation eases only once R grows beyond this separability threshold. A controlled synthetic experiment with 10 seeds confirms the predicted curves: naive balanced accuracy rises from the Bayes level toward 1.0 as recording-specific nuisance variation grows, while honest recording-grouped evaluation declines to chance, with a gap reaching about 0.5. On the public DroneRF dataset, pooled leave-one-recording-out cross-validation shows drone type identification, AR versus Bebop, collapsing from a naive macro-F1 of 0.74 to 0.46, the two-class chance level. A leakage-pathway ablation attributes essentially all of the inflation to segment-level leakage.

无人机检测数据泄漏射频感知模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。