arXiv:2608.14633eess.SPcs.LG2026-08

在罕见心律失常检测中,发现标签错误导致假阳性,影响模型可靠性。

Wolff-Parkinson-White Detection at 471:1 Class Imbalance: A Leakage-Controlled Study of the Data Bottleneck

论文配图:Wolff-Parkinson-White Detection at 471:1 Class Imbalance: A Leakage-Controlled Study of the Data Bottleneck
图 1 · 摘自论文原文
  • 构建双数据集融合模型,严格控制数据泄露,评估信号表征效果。
  • 模型最高平均精度0.595,受标签噪声和数据瓶颈限制未达饱和。
  • 揭示负样本中存在预激记录,说明标签问题来自数据本身。

Wolff-Parkinson-White (WPW) 综合征是一种先天性心脏预激,临床重要但常在静息12导联心电图中被遗漏。检测困难:特征细微且患病率极低。我们整合两个公开12导联数据集——PTB-XL与Chapman-Shaoxing-Ningbo,共66,951个记录,其中142例为WPW,患病率为0.21%(约471:1)。在预设的、严格防泄露的协议下,使用保留折叠仅接触一次,对比七种信号表示方法,固定划分与评估方式。在有限算力下,增加多样性与容量未能突破性能上限:最正交的检测器显著下降,特征融合模型表现相当于双成员投票,卷积网络未超越小波检测器,自监督预训练未通过预设阈值。无泄露学习曲线显示,在完整115个阳性样本下,最强部署检测器仍呈上升趋势(配对90%到100%差异+0.027,95% CI [0.019, 0.033]),表明尚未饱和。误差分析基于独立证据发现,漏检病例的QRS波更窄,该现象在管道外机器测量中得到验证,且其显著性依赖于所用分割算法;误标样本在漏检中无富集;部分看似假阳性实为数据集自身标注为预激的记录,提示负类标签存在缺陷。非嵌套选择带来的乐观偏差为0.11至0.13平均精度。最终输出为冻结参考分布中的百分位排名,非概率值。在保留折叠上,针对14个阳性样本,平均精度达0.595,ROC面积为0.950。该系统为筛查预过滤工具,非诊断工具。

原文摘要 · Abstract (English)

Wolff-Parkinson-White (WPW) syndrome is a congenital cardiac pre-excitation, clinically important and often missed on the resting 12-lead ECG. Detection is hard: the signature is subtle and the condition rare. We pool two public 12-lead corpora, PTB-XL and Chapman-Shaoxing-Ningbo: 66,951 recordings, 142 of them WPW, a prevalence of 0.21% (about 471:1). Under one pre-specified, leakage-controlled protocol, with a held-out fold contacted exactly once, we compare seven representations of the signal, holding the split and the evaluation fixed. Within these corpora and under a modest compute budget, added diversity and capacity do not raise the ceiling: the most orthogonal detector significantly hurts, a feature-union model matches a two-member vote, a convolutional network reaches the wavelet detector without exceeding it, and self-supervised pretraining fails a pre-specified gate. A leak-free learning curve, re-selecting features at every size, still rises at the full 115 positives for the strongest deployed detector (paired 90-to-100% difference +0.027, 95% CI [0.019, 0.033]), so it is not shown to have saturated. An error analysis tested against independent evidence finds that the missed cases have a narrower QRS, confirmed by an on-machine measurement outside our pipeline after we show the sign of this effect depends on which delineator measures it; that uncertain labels show no enrichment among the misses; and that some apparent false positives are recordings the corpus itself codes as pre-excited, placing part of the label problem in the negative class. We measure the optimism of non-nested selection at 0.11 to 0.13 average precision. The deployed output is a percentile rank in a frozen reference distribution, not a probability. On the held-out fold, on 14 positives, it reaches an average precision of 0.595 and an ROC area of 0.950. It is a screening pre-filter, not a diagnostic tool.

心电图罕见病检测数据标签模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。