针对生存分析中的区间删失数据,提出新型非参数提升方法。
Boosting methods for interval-censored data with regression and classification
- 用删失无偏变换调整损失函数与响应变量
- 在有限样本下保持高预测精度且理论性质完备
- 适合医学、工程等存在区间删失数据的领域
提升方法在机器学习与统计学领域受到广泛关注。传统提升算法针对完全观测样本设计,难以应对真实世界问题,尤其在区间删失数据场景中表现不佳。这类数据常见于生存分析与时间至事件研究中,事件发生时间未知但落在已知区间内。有效处理此类数据对医学研究、可靠性工程及社会科学至关重要。本文提出针对回归与分类任务的新型非参数提升方法,利用删失无偏变换调整损失函数并重构响应变量,同时保持模型准确性。通过函数梯度下降实现,确保可扩展性与适应性。我们严格建立了其理论性质,包括最优性与均方误差权衡。所提方法不仅为区间删失数据提供稳健的预测框架,还拓展了现有提升技术的应用范围。实证研究表明,其在多种有限样本情景下均表现出色,凸显实际应用价值。
原文摘要 · Abstract (English)
Boosting has garnered significant interest across both machine learning and statistical communities. Traditional boosting algorithms, designed for fully observed random samples, often struggle with real-world problems, particularly with interval-censored data. This type of data is common in survival analysis and time-to-event studies where exact event times are unobserved but fall within known intervals. Effective handling of such data is crucial in fields like medical research, reliability engineering, and social sciences. In this work, we introduce novel nonparametric boosting methods for regression and classification tasks with interval-censored data. Our approaches leverage censoring unbiased transformations to adjust loss functions and impute transformed responses while maintaining model accuracy. Implemented via functional gradient descent, these methods ensure scalability and adaptability. We rigorously establish their theoretical properties, including optimality and mean squared error trade-offs. Our proposed methods not only offer a robust framework for enhancing predictive accuracy in domains where interval-censored data are common but also complement existing work, expanding the applicability of existing boosting techniques. Empirical studies demonstrate robust performance across various finite-sample scenarios, highlighting the practical utility of our approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。