arXiv:2510.15363stat.MLcs.AI2025-10NeurIPS

首次为非独立同分布数据中的核回归建立理论框架,解决去噪得分学习的依赖问题。

Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning

  • 提出块分解方法,实现对依赖数据的精确集中分析
  • 推导出依赖于核谱、因果结构和采样机制的泛化误差界
  • 为去噪得分学习提供采样策略指导,适用于真实场景数据

核岭回归(KRR)是机器学习的基础工具,近期研究揭示其与神经网络的关联。然而现有理论主要针对独立同分布(i.i.d.)场景,而现实数据常具有结构化依赖关系,尤其在去噪得分学习中,多个噪声观测源自共享底层信号。本文首次系统研究具有信号-噪声因果结构的非独立同分布数据下的KRR泛化性能。通过提出新颖的分块分解方法,实现对依赖数据的精确集中分析,推导出显式依赖于:(1) 核谱,(2) 因果结构参数,(3) 采样机制(包括信号与噪声样本量相对比例)的过拟合风险界。进一步将结果应用于去噪得分学习,建立了泛化保证并提供了采样策略的合理指导。本工作推进了KRR理论,同时为现代机器学习中的依赖数据分析提供实用工具。

原文摘要 · Abstract (English)

Kernel ridge regression (KRR) is a foundational tool in machine learning, with recent work emphasizing its connections to neural networks. However, existing theory primarily addresses the i.i.d. setting, while real-world data often exhibits structured dependencies - particularly in applications like denoising score learning where multiple noisy observations derive from shared underlying signals. We present the first systematic study of KRR generalization for non-i.i.d. data with signal-noise causal structure, where observations represent different noisy views of common signals. By developing a novel blockwise decomposition method that enables precise concentration analysis for dependent data, we derive excess risk bounds for KRR that explicitly depend on: (1) the kernel spectrum, (2) causal structure parameters, and (3) sampling mechanisms (including relative sample sizes for signals and noises). We further apply our results to denoising score learning, establishing generalization guarantees and providing principled guidance for sampling noisy data points. This work advances KRR theory while providing practical tools for analyzing dependent data in modern machine learning applications.

核回归去噪得分依赖数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。