arXiv:2410.19725stat.MLcs.LG2024-10被引 8

主动采集数据可显著加速线性算子学习,优于传统随机采样。

On the Benefits of Active Data Collection in Operator Learning

  • 基于协方差核的特征值衰减速率设计主动采样策略
  • 主动策略误差收敛速度可任意快,被动策略上限为线性衰减
  • 适合需高效学习算子的场景,如科学计算与建模

研究在目标算子为线性、输入函数服从均值为零的连续协方差核随机过程时,主动数据采集策略在算子学习中的优势。通过主动采集,误差收敛速率由协方差核特征值的衰减速率决定;当特征值衰减足够快时,可实现任意快的收敛。相比之下,被动(独立同分布)采样策略的收敛速率始终不超过线性(∼n⁻¹),且对任意协方差核特征值衰减率,均存在不可消除的下界。结果表明,在此设定下,主动数据采集明显优于被动策略。

原文摘要 · Abstract (English)

We study active data collection strategies for operator learning when the target operator is linear and the input functions are drawn from a mean-zero stochastic process with continuous covariance kernels. With an active data collection strategy, we establish an error convergence rate in terms of the decay rate of the eigenvalues of the covariance kernel. We can achieve arbitrarily fast error convergence rates with sufficiently rapid eigenvalue decay of the covariance kernels. This contrasts with the passive (i.i.d.) data collection strategies, where the convergence rate is never faster than linear decay ($\sim n^{-1}$). In fact, for our setting, we show a \emph{non-vanishing} lower bound for any passive data collection strategy, regardless of the eigenvalues decay rate of the covariance kernel. Overall, our results show the benefit of active data collection strategies in operator learning over their passive counterparts.

算子学习主动学习数据采集收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。