构建统一数据模型,让机器学习更好处理复杂物理仿真。
PLAID: A Unified Data Model for Machine Learning on Heterogeneous Physics Simulations
- 提出PLAID统一数据层,支持异构物理仿真数据的高效处理
- 发布6个真实工业场景数据集,覆盖结构力学与流体动力学
- 开源工具链+社区评测平台,推动可复现的开放基准测试
基于机器学习的代理模型已成为加速模拟驱动科学工作流的强大工具,但其应用受限于缺乏大规模、多样且标准化的物理仿真数据集。现有基准多聚焦于狭窄领域或依赖简化数据模型,未能捕捉由可变几何、网格和拓扑带来的异质性,而这是评估真实场景泛化能力的关键。我们提出PLAID(Physics-Learning AI Data model),一种统一且可扩展的异构物理仿真数据层,保留仿真数据的全部复杂性,同时支持高效的机器学习工作流。配套提供数据集构建与操作库(github.com/PLAID-lib/plaid)。我们发布了六个涵盖结构力学与计算流体动力学的数据集,设计反映真实工业场景,并提供标准化基准。框架包含可复现的评估协议,并集成至Hugging Face,支持开放、社区驱动的基准测试(huggingface.co/PLAIDcompetitions)。
原文摘要 · Abstract (English)
Machine learning-based surrogate models have emerged as a powerful tool to accelerate simulation-driven scientific workflows, but their adoption is limited by the lack of large-scale, diverse, and standardized datasets for physics-based simulations. Existing benchmarks often focus on narrow domains or rely on simplified data models, and fail to capture the heterogeneity arising from variable geometries, meshes, and topologies, which is critical for assessing generalization in realistic settings. We introduce PLAID (Physics-Learning AI Data model), a unified and extensible data layer for heterogeneous physics simulations. It preserves the full complexity of simulation data while enabling efficient and scalable machine learning workflows, together with a library for dataset construction and manipulation~(\href{https://github.com/PLAID-lib/plaid}{github.com/PLAID-lib/plaid}). We release six datasets covering structural mechanics and computational fluid dynamics, designed to reflect realistic industrial scenarios and provide standardized benchmarks. The framework includes reproducible evaluation protocols and is integrated with Hugging Face to enable open, community-driven benchmarking with active user participation (\href{https://huggingface.co/PLAIDcompetitions}{huggingface.co/PLAIDcompetitions}).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。