arXiv:2510.06434cs.LGstat.ML2025-10被引 1

提出新方法在多轨迹数据中近乎最优地恢复参数,无需混合假设。

Nearly Instance-Optimal Parameter Recovery from Many Trajectories via Hellinger Localization

  • 基于赫林格距离的路径测度控制,结合轨迹费雪信息进行参数空间定位。
  • 在四种不同模型上实现接近渐近正态性的最优率,优于传统方法。
  • 适合研究序列建模、非独立数据学习的学者,尤其关注高效参数估计者。

从时间相关数据中学习是现代机器学习的核心问题。然而,对多轨迹情形(即多个独立随机过程的观测)的理解仍不充分。此类设置既反映大型基础模型的训练流程,又可在无需典型混合假设下实现学习。现有理论仅对依赖协变量的最小二乘回归给出实例最优界;对于更一般模型或损失函数,通常依赖两类简化:一是归约为i.i.d.学习(有效样本量仅随轨迹数增长),二是利用单轨迹混合结果(有效样本量受混合时间压缩)。本文通过赫林格局部化框架,显著拓展了多轨迹下的实例最优率适用范围。该方法先在路径测度层面通过归约控制平方赫林格距离,再在参数空间以轨迹费雪信息加权进行局部化,从而在广泛条件下实现与总数据预算相关的实例最优界。我们在四个案例中验证:简单马尔可夫链混合模型、非高斯噪声下的依赖线性回归、具非单调激活的广义线性模型,以及线性注意力序列模型。所有情况下,边界均接近渐近正态性下的最优率,显著优于标准归约方法。

原文摘要 · Abstract (English)

Learning from temporally-correlated data is a core facet of modern machine learning. Yet our understanding of sequential learning remains incomplete, particularly in the multi-trajectory setting where data consists of many independent realizations of a time-indexed stochastic process. This important regime both reflects modern training pipelines such as for large foundation models, and offers the potential for learning without the typical mixing assumptions made in the single-trajectory case. However, instance-optimal bounds are known only for least-squares regression with dependent covariates; for more general models or loss functions, the only broadly applicable guarantees result from a reduction to either i.i.d. learning, with effective sample size scaling only in the number of trajectories, or an existing single-trajectory result when each individual trajectory mixes, with effective sample size scaling as the full data budget deflated by the mixing-time. In this work, we significantly broaden the scope of instance-optimal rates in multi-trajectory settings via the Hellinger localization framework, a general approach for maximum likelihood estimation. Our method proceeds by first controlling the squared Hellinger distance at the path-measure level via a reduction to i.i.d. learning, followed by localization as a quadratic form in parameter space weighted by the trajectory Fisher information. This yields instance-optimal bounds that scale with the full data budget under a broad set of conditions. We instantiate our framework across four diverse case studies: a simple mixture of Markov chains, dependent linear regression under non-Gaussian noise, generalized linear models with non-monotonic activations, and linear-attention sequence models. In all cases, our bounds nearly match the instance-optimal rates from asymptotic normality, substantially improving over standard reductions.

参数估计序列学习概率建模优化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。