arXiv:2512.20941cs.LGcs.AI2025-12

提出双三角翼多保真度数据集,揭示神经网络预测精度与数据量的规律。

A Multi-fidelity Double-Delta Wing Dataset and Empirical Scaling Laws for GNN-based Aerodynamic Field Surrogate

  • 构建双三角翼多保真度气动场数据集,含2448个流场快照
  • 发现测试误差随数据量呈幂律下降,指数为-0.6122
  • 建议每维约8个采样点,适配大模型与高效设计优化

数据驱动的代理模型正被广泛用于加速车辆设计。然而,开源的多保真度数据集以及数据量与模型性能之间关系的实证指导仍较匮乏。本研究探究了基于图神经网络(GNN)的气动场代理模型中训练数据量与预测精度的关系。我们发布了开源的双三角翼多保真度气动数据集,涵盖272种几何构型在攻角11°至19°、马赫数Ma=0.3条件下的2448个流场快照,使用涡格法(VLM)和雷诺平均N-S方程(RANS)求解器获得。几何构型采用嵌套Saltelli采样方案生成,支持未来扩展与方差敏感性分析。基于该数据集,对多保真度涡格网络(MF-VortexNet)进行了初步实证尺度研究,构建了从40到1280个快照的六组训练集,在固定训练预算下训练参数量为0.1至2.4百万的模型。结果表明,测试误差随数据量按幂律下降,指数为-0.6122,表明数据利用效率高。据此估算,最优采样密度约为每维8个样本。结果还显示,大模型具有更高的数据利用效率,暗示数据生成成本与训练预算之间存在潜在权衡。

原文摘要 · Abstract (English)

Data-driven surrogate models are increasingly adopted to accelerate vehicle design. However, open-source multi-fidelity datasets and empirical guidelines linking dataset size to model performance remain limited. This study investigates the relationship between training data size and prediction accuracy for a graph neural network (GNN) based surrogate model for aerodynamic field prediction. We release an open-source, multi-fidelity aerodynamic dataset for double-delta wings, comprising 2448 flow snapshots across 272 geometries evaluated at angles of attack from 11 (degree) to 19 (degree) at Ma=0.3 using both Vortex Lattice Method (VLM) and Reynolds-Averaged Navier-Stokes (RANS) solvers. The geometries are generated using a nested Saltelli sampling scheme to support future dataset expansion and variance-based sensitivity analysis. Using this dataset, we conduct a preliminary empirical scaling study of the MF-VortexNet surrogate by constructing six training datasets with sizes ranging from 40 to 1280 snapshots and training models with 0.1 to 2.4 million parameters under a fixed training budget. We find that the test error decreases with data size with a power-law exponent of -0.6122, indicating efficient data utilization. Based on this scaling law, we estimate that the optimal sampling density is approximately eight samples per dimension in a d-dimensional design space. The results also suggest improved data utilization efficiency for larger surrogate models, implying a potential trade-off between dataset generation cost and model training budget.

气动代理图神经网络多保真度数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。