统一评测工程设计中生成与预测模型,打破学术数据孤岛。
PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

- 用统一流程评估生成与预测模型,覆盖七类任务和28种配置。
- 在小数据到超大规模下测试,66个模型在9个工业级数据集上排名。
- 引入去偏排名机制,揭示模型性能随数据量变化的动态规律。
生成式与预测式人工智能模型在工程设计与仿真中广泛用于生成几何结构并预测物理场与标量值。然而,这些模型通常在孤立的学术数据集上评估,使用不一致的指标与流程。我们提出PhysicsBench,一个统一的基准与排行榜,对生成与预测模型采用标准化评估流程。PhysicsBench涵盖一维、二维、三维领域的七类生成与预测任务,在九个数据集上评估66个模型,包含工业级CAD/CFD/FEA仿真与公开参考数据,扩展为28种配置。统一流程与排名适用于两类模型,每类独立评分。评估覆盖从S到XL的真实有限数据规模,而非学术基准中无限训练集。通用指标套件包含几何保真度(分布距离)、物理场与标量精度,以及工程相关的场与形状有效性。BenchRank通过头对头优势图进行PageRank去偏排名,所有质量指标均被排序,计算成本另设效率视图。跨任务分析显示,大型模型在学术数据上的表现弱预测其在小数据下的排名;在七项任务中,有六项的最优模型随数据规模变化,无模型在超过一项任务中持续领先。PhysicsBench将‘最先进’从自我宣称转化为公开可查的模型选择基础。
原文摘要 · Abstract (English)
Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. Yet these models are typically evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. We present PhysicsBench, a unified benchmark and leaderboard that evaluates generative and predictive models under one standardized procedure. PhysicsBench spans seven generation and prediction tasks across 1D, 2D, and 3D domains and ranks 66 models on nine datasets, comprising industrial-scale CAD/CFD/FEA simulations and public references, expanded into 28 configurations. One procedure and ranking apply to both families, each ranked within its own tasks. Evaluation spans realistic, limited data scales from S to XL rather than the unlimited training sets common in academic benchmarks. A common metric suite captures geometric fidelity with distributional distances, physical-field and scalar accuracy, and engineering-specific field- and shape-validity. BenchRank debiases correlated metrics and ranks by PageRank over a head-to-head dominance graph, so every reported quality metric is also ranked, with computational cost in a separate efficiency view. Across tasks, an architecture's large-scale academic standing weakly predicts its small-data ranking. The top model changes with data scale in six of the seven tasks, and no model leads more than one task. PhysicsBench turns "state-of-the-art" from a self-reported claim into an openly published foundation for model selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。