提出盲点质量度量,量化模型在罕见部署状态下的风险。
Blind-Spot Mass: A Good-Turing Framework for Quantifying Deployment Coverage Risk in Machine Learning Systems
- 用古德-图林方法估算低频状态的总概率质量
- 发现临床与可穿戴数据中盲点质量在阈值5时均达95%
- 帮助识别高风险领域,指导数据补充与模型约束
盲点质量是一种基于古德-图林框架的机器学习系统部署覆盖风险量化方法。现代机器学习系统的运行状态分布常呈长尾特征,导致大量有效但罕见的状态在有限训练与评估数据中缺乏支持,形成‘覆盖盲区’:模型在标准测试集上表现准确,却在部署状态空间的大片区域不可靠。本文提出盲点质量 B_n(tau),用于估计经验支持低于阈值 tau 的状态所占总概率质量。该指标基于古德-图林未见物种估计法计算,能合理预估部署分布中处于可靠性关键但支持不足区域的比例。进一步推导出覆盖限制下的准确率上限,将整体性能分解为受支持与盲点两部分,分离能力极限与数据局限。在可穿戴人体活动识别(HAR)中验证该框架,并在包含275例入院记录的MIMIC-IV数据库中复现分析,不同模态、特征空间、标签空间和应用场景下,盲点质量曲线在 tau=5 时均收敛至95%。这一跨域一致性表明,盲点质量是通用的组合覆盖风险量化方法,非特定应用结果。其分解机制可定位高风险活动或临床状态,为工业界提供数据采集、归一化及物理/领域约束的可操作指导。
原文摘要 · Abstract (English)
Blind-spot mass is a Good-Turing framework for quantifying deployment coverage risk in machine learning. In modern ML systems, operational state distributions are often heavy-tailed, implying that a long tail of valid but rare states is structurally under-supported in finite training and evaluation data. This creates a form of 'coverage blindness': models can appear accurate on standard test sets yet remain unreliable across large regions of the deployment state space. We propose blind-spot mass B_n(tau), a deployment metric estimating the total probability mass assigned to states whose empirical support falls below a threshold tau. B_n(tau) is computed using Good-Turing unseen-species estimation and yields a principled estimate of how much of the operational distribution lies in reliability-critical, under-supported regimes. We further derive a coverage-imposed accuracy ceiling, decomposing overall performance into supported and blind components and separating capacity limits from data limits. We validate the framework in wearable human activity recognition (HAR) using wrist-worn inertial data. We then replicate the same analysis in the MIMIC-IV hospital database with 275 admissions, where the blind-spot mass curve converges to the same 95% at tau = 5 across clinical state abstractions. This replication across structurally independent domains - differing in modality, feature space, label space, and application - shows that blind-spot mass is a general ML methodology for quantifying combinatorial coverage risk, not an application-specific artifact. Blind-spot decomposition identifies which activities or clinical regimes dominate risk, providing actionable guidance for industrial practitioners on targeted data collection, normalization/renormalization, and physics- or domain-informed constraints for safer deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。