用合成数据模拟物理方程结构,提升机器学习泛化能力。
Synthics: Synthetic Physics-like Datasets for Machine Learning

- 基于贝叶斯语法生成符合物理规律的新方程。
- 合成数据在8个结构特征上匹配真实数据,优于传统方法。
- 适合需要高质量训练数据的物理建模与自动化机器学习研究者。
代表性数据对机器学习至关重要,但真实数据采集常受限。合成数据是可行方案,前提是其结构需忠实反映真实观测。本文提出一种生成结构类比物理方程的合成回归数据的方法:利用贝叶斯概率上下文无关语法捕捉方程语料库的代数结构,并从中采样新方程。为确保输入处于物理可实现域,通过非侵入探测刻画每条方程的适用域,并恢复变量间约束。输入采样模拟真实实验条件,采用混合均匀分布与截断正态分布从有效域的随机子区间中采样。通过柯尔莫哥洛夫-斯米尔诺夫检验,合成数据在费曼方程语料库的八个结构特征上均匹配,而未平滑的纯概率语法仅匹配两项,证明贝叶斯先验对结构保真性至关重要。下游超参数调优任务中,基于合成数据训练的梯度提升回归器在20个配置中平均选中第6名,与真实数据调优效果一致,显著优于随机表达式树(第10名)和噪声(第19名)。
原文摘要 · Abstract (English)
Representative data is fundamental in machine learning, as limited data hinders generalisation. Collecting sufficient real-world samples is often infeasible. Synthetic data generation offers a practical solution, but only if the generated data faithfully reflects the structure of real observations. In this paper, a method for generating synthetic regression datasets that structurally resemble physics equations from a given equation corpus is presented. The approach uses a Bayesian Probabilistic Context-Free Grammar to capture the underlying algebraic structure of the corpus, from which novel equations are sampled. To ensure the generated inputs lie within a physically meaningful domain, the applicability domain is characterised for each equation through non-intrusive probing, also recovering inter-variable constraints. Input sampling further mimics realistic experimental conditions by drawing from random sub-ranges of the valid domain with mixed uniform and truncated normal distributions. The generated data is statistically validated against the Feynman equation corpus using Kolmogorov-Smirnov tests. The generated equations match the corpus on all of the eight studied structural features, compared to only two for an unsmoothed purely probabilistic grammar, demonstrating that the Bayesian prior is essential for structural fidelity given the size of the corpus. In a downstream hyperparameter-tuning task, a gradient-boosted regressor tuned on the synthetic data picks, on average, the 6th-best configuration out of 20 on real data, matching the result of tuning on real data itself and substantially outperforming random expression trees (10th) and noise (19th).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。