arXiv:2507.02215stat.MLcs.LG2025-07被引 1

新方法在严重噪声下仍能高效逼近函数,适合高噪声数据建模。

Hybrid least squares for learning functions from highly noisy data

  • 融合克里斯托弗尔采样与最优实验设计,自适应生成样本点。
  • 在噪声极大时样本复杂度更低,计算效率提升显著。
  • 适用于随机场期望值建模,尤其适合金融模拟等复杂场景。

针对严重污染数据下的条件期望估计需求,本文研究了带噪声的最小二乘函数逼近问题。现有小噪声场景有效的方法在大噪声下表现不佳。为此,提出一种结合克里斯托弗尔采样与最优实验设计的混合方法。该算法在样本点生成和噪声平滑方面均具备良好优化性,相较于已有方法,在计算效率和样本复杂度上均有提升。进一步将算法扩展至凸性约束情形,并保持相同理论保证。当目标函数为随机场期望时,引入自适应随机子空间,证明了自适应过程的逼近能力。理论结果在合成数据及计算金融中的复杂随机模拟问题上得到数值验证。

原文摘要 · Abstract (English)

Motivated by the need for efficient estimation of conditional expectations, we consider a least-squares function approximation problem with heavily polluted data. Existing methods that are effective in the small-noise regime are suboptimal when large noise is present. To address this issue, we propose a hybrid approach that combines Christoffel sampling with optimal experimental design. We show that the proposed algorithm enjoys appropriate optimality properties for both sample point generation and noise mollification, leading to improved computational efficiency and sample complexity compared to existing methods. We also extend the algorithm to convexity-constrained settings with similar theoretical guarantees. When the target function is defined as the expectation of a random field, we further extend our approach to leverage adaptive random subspaces and establish results on the approximation capacity of the adaptive procedure. Our theoretical findings are supported by numerical studies on both synthetic data and on a more challenging stochastic simulation problem in computational finance.

函数逼近高噪声最优设计随机模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。