针对极端值采样数据,提出新型半监督分类方法,显著提升罕见事件检测效果。
Fractionally Supervised Classification with Maxima Nominated Samples

- 引入潜在变量建模极值样本与剩余单元的类别关系
- 在罕见事件混合模型中表现优于忽略排名信息的旧方法
- 适用于筛查、环境监测等极端值主导的场景
分数监督分类(FSC)提供了一种结合已标记与未标记数据的灵活框架,但现有方法假设为简单随机采样。在许多实际应用中,保留的观测是某组中的极值而非随机个体,尤其在目标群体稀有时,最大值提名采样(NS)可增强样本信息量,如筛查、环境监测、重复测试和可靠性研究。在此类设计下,似然函数发生根本性变化,传统FSC的EM算法不再适用。本文通过引入潜变量表示所观察极大值的类别及组内其余单元的隐含构成,构建了适用于提名样本的正确FSC方法,得到有效的EM算法和加权似然程序。方法以通用形式呈现,并应用于罕见事件污染的正态混合模型;模拟显示其显著优于忽略额外排序信息的错误设定;真实数据分析验证了其实际价值。
原文摘要 · Abstract (English)
Fractionally supervised classification (FSC) offers a flexible framework for combining labeled and unlabeled data in model-based classification, but existing formulations assume simple random sampling. In many applications, however, the retained observation is an extreme order statistic from a set rather than a randomly selected unit. This is particularly appealing when the target population is rare, since maxima nomination sampling (NS) can enrich the sample with the most informative observations, as in screening, environmental monitoring, repeated testing, and reliability studies. Under such designs, the likelihood function changes fundamentally, and the usual FSC EM construction is no longer valid. We develop FSC for nominated samples by introducing a latent representation that accounts for both the class membership of the observed maximum and the latent composition of the remaining units in the set. The resulting method yields a proper EM algorithm and a coherent weighted-likelihood FSC procedure for NS data. We present the methodology in general form, illustrate it for a rare-event contamination normal mixtures, and show through simulation that it substantially improves on the misspecified alternative by ignoring the extra rank information of such data. A real-data analysis demonstrates its practical value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。