在生存数据去删失中,用预算控制标签获取,提升模型性能。
Budget-constrained Active Learning to Effectively De-censor Survival Data
- 基于预算的主动学习,可部分解封删失数据以获取更多信息。
- 实验表明,在多个生存任务上优于其他方法,且计算效率与主流方法相当。
- 适合医疗健康领域中数据标注成本高、存在删失的场景。
标准监督学习从标记数据集中学习模型。给定少量标记样本和大量未标记样本,有预算的学习者可使用预算支付以获取部分未标记样本的标签,从而构建模型。本文探讨在包含(右)删失样本的生存数据中应用预算学习,其中仅知实例事件时间的下限。学习者可通过付费部分解封删失实例——例如将(3年,删失)变为(7.2年,非删失),或获得更详细信息(如变为(4年,删失))。这模拟了真实世界随访中无法完全解除删失的情况,信息量取决于预算和数据特性。我们提供了实验和理论结果,展示如何将先进预算学习算法应用于生存数据,并揭示其局限性。所提方法在边界和时间复杂度上与标准主动学习方法BatchBALD渐近等价。多个生存任务的实证分析显示,该模型在多个基准上表现优于现有方法。
原文摘要 · Abstract (English)
Standard supervised learners attempt to learn a model from a labeled dataset. Given a small set of labeled instances, and a pool of unlabeled instances, a budgeted learner can use its given budget to pay to acquire the labels of some unlabeled instances, which it can then use to produce a model. Here, we explore budgeted learning in the context of survival datasets, which include (right) censored instances, where we know only a lower bound on an instance's time-to-event. Here, that learner can pay to (partially) label a censored instance -- e.g., to acquire the actual time for an instance [perhaps go from (3 yr, censored) to (7.2 yr, uncensored)], or other variants [e.g., learn about one more year, so go from (3 yr, censored) to either (4 yr, censored) or perhaps (3.2 yr, uncensored)]. This serves as a model of real world data collection, where follow-up with censored patients does not always lead to uncensoring, and how much information is given to the learner model during data collection is a function of the budget and the nature of the data itself. We provide both experimental and theoretical results for how to apply state-of-the-art budgeted learning algorithms to survival data and the respective limitations that exist in doing so. Our approach provides bounds and time complexity asymptotically equivalent to the standard active learning method BatchBALD. Moreover, empirical analysis on several survival tasks show that our model performs better than other potential approaches on several benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。