arXiv:2608.11019cs.LG2026-08

用频域采样生成高效物理数据,显著降低训练需求。

DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling

论文配图:DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling
图 1 · 摘自论文原文
  • 通过逆傅里叶变换控制关键频谱分量生成合理数据
  • 在多个系统中减少40%数据量,准确率损失<2%
  • 适合电池等复杂系统的低样本建模,且可跨化学体系迁移

建模由偏微分方程(PDE)描述的时空动力系统面临两大挑战:要么依赖昂贵的物理模拟器,要么需要大量训练数据,而纯数据驱动模型在下游动态条件下泛化能力差。本文提出DEFT,一种频域数据采样方法,通过识别物理系统的主导傅里叶模式,并系统调节其振幅与相位,利用逆离散傅里叶变换生成物理一致的训练数据。同时推导了该方法的泛化界,提供选择K的理论依据。通过三组实验验证:第一,在经典PDE求解中,当系统由少数显著频率成分主导时,优于传统方法;第二,在PDEBench的扩散-吸附与Burgers方程上作为数据价值过滤器,使数据需求减少40%,预测准确率损失小于2%;第三,在电池退化PDE系统中,跨多种测试数据集保持高精度,R²值超过0.99,且学习到的频域特征仅需20%微调数据即可迁移至其他电池化学体系。结果表明,DEFT是高效的算子学习数据采样方法。

原文摘要 · Abstract (English)

Modeling spatiotemporal dynamical systems governed by partial differential equations (PDEs) poses two major challenges: it either requires expensive physics-based simulators that entail iterative numerical solving at high computational cost, or it depends on abundant training data, yet purely data-driven models often generalize poorly to downstream dynamic operating conditions. We propose DEFT, a frequency-domain data sampling method that identifies the dominant Fourier modes of a physical system and systematically varies the corresponding amplitudes and phases to generate physically consistent training data via the inverse discrete Fourier transform. In addition, we derive a generalization bound of this method. We note that it also provides a theoretically principled criterion for selecting $K$. We evaluate the proposed method through three sets of experiments, each targeting a distinct aspect of its utility. First, we validate the framework on canonical PDEs solving demonstrating that it outperforms traditional methods when the system is dominated by a few prominent frequency components. Second, we employ DEFT as a data-value filter on the diffusion--sorption and Burgers equations of PDEBench, showing that it reduces data requirements by $40\%$ while sacrificing less than $2\%$ in predictive accuracy. Third, to evaluate DEFT for more challenging and practically relevant problems, we validate it in the battery degradation PDE system, achieving consistently high predictive accuracy across various test datasets with $R^2$ values exceeding $0.99$. Moreover, the learned frequency-domain features transfer to other battery chemistries with only $20\%$ of the fine-tuning data. These results demonstrate that DEFT is an effective data-sampling method for efficient operator learning.

频域建模低样本学习物理信息神经网络数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。