构建通用任务框架,加速科学工程中AI模型的评估与进步
Accelerating scientific discovery with the common task framework
- 设计统一挑战数据集,覆盖预测、状态重建等常见科学目标
- 支持在小样本和噪声数据下对比不同AI算法性能
- 适合从事科学计算与跨领域建模的研究者参考
机器学习与人工智能正在重塑工程、物理和生命科学中动态系统的表征与控制。这些新兴建模范式需要可比的评估指标,以应对多样化的科学目标,包括预测、状态重构、泛化能力和控制,并兼顾数据有限与测量噪声等现实挑战。本文提出科学与工程领域的通用任务框架(CTF),包含不断扩展的挑战数据集,涵盖多种实际且常见的科学目标。该框架是推动机器学习/人工智能算法在语音识别、自然语言处理和计算机视觉等传统领域快速发展的关键使能技术。当前亟需一套客观的评估指标体系,以衡量各类算法在科学与工程实践中迅速发展和部署的表现。
原文摘要 · Abstract (English)
Machine learning (ML) and artificial intelligence (AI) algorithms are transforming and empowering the characterization and control of dynamic systems in the engineering, physical, and biological sciences. These emerging modeling paradigms require comparative metrics to evaluate a diverse set of scientific objectives, including forecasting, state reconstruction, generalization, and control, while also considering limited data scenarios and noisy measurements. We introduce a common task framework (CTF) for science and engineering, which features a growing collection of challenge data sets with a diverse set of practical and common objectives. The CTF is a critically enabling technology that has contributed to the rapid advance of ML/AI algorithms in traditional applications such as speech recognition, language processing, and computer vision. There is a critical need for the objective metrics of a CTF to compare the diverse algorithms being rapidly developed and deployed in practice today across science and engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。