解决物理仿真训练中的数据分布偏差问题,提升模型精度与泛化能力。
M$^3$: Reframing Training Measures for Discretized Physical Simulations

- 按物理变化分区并多尺度分配监督信号,平衡训练数据分布。
- 在大规模体积场景中误差降低4.7倍,子采样下仍优于高分辨率数据。
- 适合工业级物理模拟任务,尤其适用于数据稀缺或分布不均场景。
物理仿真中的神经代理模型在连续域的离散采样上训练,导致经验分布不均,引发监督偏差并造成物理保真度的空间不一致。为缓解这种由度量引起的偏差,我们提出M$^3$(多尺度莫顿度量)框架,通过根据物理变化划分空间并在多尺度上分配监督信号,实现训练度量的平衡。在三个具有不同离散化的工业级数据集上应用,M$^3$在连续物理域中持续提升预测性能,最大体积案例中误差降低4.7倍。在极端子采样(160M → 16M → 1.6M点)下,M$^3$训练模型仍优于高分辨率数据训练的模型,物理加权相对$L_2$误差降低3–4倍,对应均方误差最高降低13倍。结果表明数据分布是算子学习的关键因素,M$^3$为物理一致性建模提供了可扩展、数据高效的解决方案。
原文摘要 · Abstract (English)
Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical measure leads to uneven supervision, biasing optimization and causing spatial inconsistencies in physical fidelity. To mitigate this measure-induced bias, we propose M$^3$ (Multi-scale Morton Measure), a scalable framework that balances training measures by partitioning space according to physical variation and allocating supervision across multiple scales. Applied to three industrial-scale datasets with diverse discretizations, M$^3$ consistently improves predictions in the continuous physical domain, achieving up to 4.7$\times$ lower error in large-scale volumetric cases. These gains persist under aggressive subsampling (160M $\rightarrow$ 16M $\rightarrow$ 1.6M points), where M$^3$-trained models outperform those trained on higher-resolution data, reducing physics-weighted relative $L_2$ error by 3--4$\times$ and the corresponding MSE by up to 13$\times$. These results highlight data distribution as a key factor in operator learning and position M$^3$ as a scalable, data-efficient approach for physically consistent modeling. Code is available at https://github.com/PhysDataRefine/M3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。