构建科学机器学习统一基准体系,打通跨领域研究壁垒。
An MLCommons Scientific Benchmarks Ontology
- 基于社区协作建立科学基准本体,整合物理、化学等多领域资源
- 提出六维评分标准,支持新基准提交与质量评估
- 为科研人员和开发者提供可复现、跨领域的基准选择框架
科学机器学习研究涵盖广泛领域与数据模态,但现有基准体系分散且缺乏统一标准,阻碍了机器学习在关键科学场景中的创新应用。本文通过协同社区努力,扩展MLCommons生态,提出覆盖物理、化学、材料科学、生物学、气候科学等领域的科学基准本体。该工作整合了XAI-BENCH、FastML Science Benchmarks、PDEBench及SciMLBench等前期成果,构建了涵盖科学、应用与系统级的统一分类体系。新基准可通过由MLCommons科学工作组协调的开放提交流程加入,并依据六类评分标准评估,确保高质量。架构具备可扩展性,支持未来新兴科学与AI/ML模式的识别。本体为科学机器学习的可复现、跨领域基准测试提供了标准化基础。配套网页将持续更新:https://mlcommons-science.github.io/benchmark/
原文摘要 · Abstract (English)
Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。