提出新准则,解决强相关变量下的模型选择难题
A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs
- 用编码长度正则化海量模型,统一复杂度衡量标准
- 在弱假设下保证选型一致,且模型误设时仍有效
- 适合高维强相关数据,尤其适合模型类不确定场景
当预测变量高度相关且模型类别不确定时,模型选择面临巨大挑战,尤其当候选模型数量呈指数级增长。本文提出描述复杂度信息准则(DCIC),通过满足Kraft条件的码长对大规模候选模型集合进行正则化。在子韦伯尔噪声条件下,无需依赖RIP型条件,即可通过近似误差分离实现选择一致性,并获得非渐近的极小风险界,该界在模型误设时依然成立。相同的编码原理可将异质模型类映射到统一复杂度尺度,仅需较小的类别识别代价,从而实现类别-模型联合恢复及跨类风险自适应。进一步设计了基于复杂度引导的搜索路径,显式揭示计算与统计间的权衡:大惩罚下保留区域为多项式规模且高概率成立;小惩罚则逼近最优风险基准。数值实验表明,在强相关和模型类不确定性下仍具稳定支撑恢复能力与优异估计性能。
原文摘要 · Abstract (English)
Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under sub-Weibull noise, we establish selection consistency through approximation-error separation without relying on RIP-type conditions, together with nonasymptotic oracle risk bounds that remain valid under model misspecification. The same coding principle places heterogeneous classes on a common complexity scale at a small additional class-identification cost. This extension yields class--model recovery under suitable identifiability conditions and risk adaptation across classes. We further develop a complexity-guided search path that makes the computation--statistics trade-off explicit. Large penalties yield polynomial-size retained search regions with high probability, whereas smaller penalties sharpen the oracle risk benchmark. Numerical experiments illustrate stable support recovery and favorable estimation performance under strong dependence and model-class uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。