材料科学中机器学习评估需警惕虚假进步,应透明化测量设计。
Lessons from the trenches on evaluating machine-learning systems in materials science
- 以统计测量理论为框架,系统分析机器学习评估中的常见问题。
- 发现评估漏洞易导致伪进展,影响研究方向与科学进步。
- 提出评估卡片机制,推动评估过程透明化与可复现。
衡量是科学知识创造的基础,确保成果可共享并支撑科学发现。随着机器学习在科学领域日益普及,如何有效评估这些系统成为保障可靠进展的关键。本文综述了机器学习在科学中评估框架的现状与未来方向,以材料科学为背景,基于统计测量理论构建通用评估框架。我们识别出跨领域的核心挑战:构念效度不足、数据质量问题、指标设计局限及基准维护困难,这些问题可能导致评估无法反映真实性能,从而引发‘伪进展’。通过分析传统基准与新兴评估方法,揭示评估选择不仅影响测量结果,更塑造研究优先级与科学演进路径。因此,亟需提升评估设计的透明度,我们提出‘评估卡片’作为结构化记录测量选择与局限的方法。本研究强调材料科学中需发展更丰富的评估工具箱,其洞见亦适用于其他面临相似挑战的科学领域。
原文摘要 · Abstract (English)
Measurements are fundamental to knowledge creation in science, enabling consistent sharing of findings and serving as the foundation for scientific discovery. As machine learning systems increasingly transform scientific fields, the question of how to effectively evaluate these systems becomes crucial for ensuring reliable progress. In this review, we examine the current state and future directions of evaluation frameworks for machine learning in science. We organize the review around a broadly applicable framework for evaluating machine learning systems through the lens of statistical measurement theory, using materials science as our primary context for examples and case studies. We identify key challenges common across machine learning evaluation such as construct validity, data quality issues, metric design limitations, and benchmark maintenance problems that can lead to phantom progress when evaluation frameworks fail to capture real-world performance needs. By examining both traditional benchmarks and emerging evaluation approaches, we demonstrate how evaluation choices fundamentally shape not only our measurements but also research priorities and scientific progress. These findings reveal the critical need for transparency in evaluation design and reporting, leading us to propose evaluation cards as a structured approach to documenting measurement choices and limitations. Our work highlights the importance of developing a more diverse toolbox of evaluation techniques for machine learning in materials science, while offering insights that can inform evaluation practices in other scientific domains where similar challenges exist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。