arXiv:2505.21626cs.LGmath.OC2025-05被引 2

优化训练数据分布,让科学机器学习模型在部署时更准更稳。

Learning where to learn: Training data distribution optimization for scientific machine learning

  • 通过双层优化设计训练数据分布,提升模型泛化能力。
  • 新方法在多种边界条件下平均误差降低,样本效率更高。
  • 适合需高鲁棒性的科学计算、微分方程求解场景。

在科学机器学习中,模型常部署于与训练参数或边界条件差异较大的场景。本文研究了‘学习何处学习’问题,旨在设计一种训练数据分布,以最小化在各类部署环境下预测的平均误差。理论分析揭示了训练分布对部署精度的影响机制。由此提出基于双层或交替优化的概率测度空间算法。离散化实现采用参数化分布类或非参数粒子梯度流,生成的优化训练分布显著优于传统非自适应设计。训练完成后,模型展现出更强的样本效率和对分布偏移的鲁棒性。该框架为偏微分方程解算子及函数学习的可论证数据采集提供了新路径。

原文摘要 · Abstract (English)

In scientific machine learning, models are routinely deployed with parameter values or boundary conditions far from those used in training. This paper studies the learning-where-to-learn problem of designing a training data distribution that minimizes average prediction error across a family of deployment regimes. A theoretical analysis shows how the training distribution shapes deployment accuracy. This motivates two adaptive algorithms based on bilevel or alternating optimization in the space of probability measures. Discretized implementations using parametric distribution classes or nonparametric particle-based gradient flows deliver optimized training distributions that outperform nonadaptive designs. Once trained, the resulting models exhibit improved sample complexity and robustness to distribution shift. This framework unlocks the potential of principled data acquisition for learning functions and solution operators of partial differential equations.

科学机器学习分布优化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。