提出高效变量选择方法,处理生物医学中的复杂数据。
Variable Selection Methods for Multivariate, Functional, and Complex Biomedical Data in the AI Age
- 基于最优子集优化,统一处理多变量、函数型等复杂数据
- 在准确率和速度上均显著优于现有方法,提速数个数量级
- 适合生物统计、人工智能等领域研究者解决高维数据筛选问题
个性化医疗与数字健康中的许多问题依赖于连续时间函数型生物标志物及其他从高分辨率患者监测中产生的复杂数据结构的分析。本文提出一种基于优化的变量选择方法,适用于度量空间中的多变量、函数型乃至更一般结果的最优子集选择。该框架可应用于线性、分位数或非参数加性模型等多种回归模型,并支持包括标量、多变量欧氏数据、函数型数据乃至随机图在内的广泛随机响应类型。分析表明,所提方法在准确率和速度上均显著优于当前先进方法,尤其在数学函数类响应场景下实现多个数量级的速度提升。尽管框架具有通用性且不针对特定科学问题,但文章自洽,聚焦生物医学应用,为生物统计、统计学及人工智能领域的专业人士提供了应对人工智能时代变量选择挑战的重要工具。
原文摘要 · Abstract (English)
Many problems within personalized medicine and digital health rely on the analysis of continuous-time functional biomarkers and other complex data structures emerging from high-resolution patient monitoring. In this context, this work proposes new optimization-based variable selection methods for multivariate, functional, and even more general outcomes in metrics spaces based on best-subset selection. Our framework applies to several types of regression models, including linear, quantile, or non parametric additive models, and to a broad range of random responses, such as univariate, multivariate Euclidean data, functional, and even random graphs. Our analysis demonstrates that our proposed methodology outperforms state-of-the-art methods in accuracy and, especially, in speed-achieving several orders of magnitude improvement over competitors across various type of statistical responses as the case of mathematical functions. While our framework is general and is not designed for a specific regression and scientific problem, the article is self-contained and focuses on biomedical applications. In the clinical areas, serves as a valuable resource for professionals in biostatistics, statistics, and artificial intelligence interested in variable selection problem in this new technological AI-era.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。