提出新型分池岭估计法,解决函数型回归在稀疏到密集采样下的最优预测问题。
Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

- 融合分池策略与核空间方法,利用所有受试者离散观测数据无偏估计算子
- 在任意采样方案下实现预测风险的极小极大最优,稀疏至密集设计均适用
- 揭示采样频率影响:标量对函数回归有1次相变,函数对函数回归最多3次
函数数据分析是将数据视为随机函数的重要统计领域。实际中,这些随机函数往往未被完整观测,仅在离散时间点测量。尽管均值和协方差估计等简化问题已在离散观测数据上得到广泛研究,但该类型数据的线性回归最优估计问题已困扰学界逾二十年。为攻克这一基础挑战,本文提出一种新方法——分池岭估计,通过结合分池策略与基于再生核希尔伯特空间(RKHS)的方法,利用所有受试者离散观测数据无偏估计算子,构建统一估计框架。该框架首次实现从稀疏到密集采样设计下,标量对函数及函数对函数回归模型的预测风险极小极大最优。方法论与理论进展准确揭示了离散采样的影响:标量对函数回归中出现一次相变,区分两种不同收敛行为;函数对函数回归中,根据预测变量与响应变量的采样频率,最多可出现三次相变。模拟实验与两个真实数据案例为所提方法提供了实证支持。
原文摘要 · Abstract (English)
Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times. While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal estimation of linear regression for this data type has remained unsolved for over two decades. To tackle this fundamental challenge, we propose a novel approach, referred to as pooling ridge estimation, which combines the advantages of pooling strategy and RKHS-based method by incorporating the unbiased estimation of operators based on discretely observed measurements from all subjects. This unified estimation framework enables us to achieve minimax optimality in prediction risk in arbitrary sampling schemes ranging from sparse to dense designs, for both scalar-on-function and function-on-function regression models. Such methodological and theoretical advances are obtained for the first time and accurately reveal the influence of discrete sampling. For scalar-on-function regression, the phase transition occurs once, separating the convergence behavior into two distinct regimes. Remarkably, for function-on-function regression, up to three phase transitions may occur, determined by the sampling frequencies of the predictor/response functions. Finally, simulation experiments and two real data examples provide empirical support for the proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。