构建跨科学领域的异常检测数据集,推动机器学习发现未知现象。
Building Machine Learning Challenges for Anomaly Detection in Science
- 设计三个跨领域科学数据集:天体物理、基因组学、极地科学。
- 提出可复用的异常检测挑战框架,支持大规模计算实验。
- 适合从事科学发现与机器学习结合的研究者使用。
科学发现常源于发现未被现有科学规律预测的模式或异常对象。这些不符合常规的事件或物体可能表明当前科学规则不完整,需引入新机制解释。但识别异常极具挑战,需全面掌握已知科学行为并将其投射到数据中以发现偏离。机器学习面临更大困难:模型不仅要精准理解科学数据,还需识别出超出训练范围的不一致数据。本文提出三个面向天体物理、基因组学和极地科学的异常检测数据集,构建可发现、可访问、可互操作、可重用(FAIR)的机器学习挑战体系。所提方法具备可扩展性,支持未来更大规模、更高算力的挑战,有望推动科学发现。
原文摘要 · Abstract (English)
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not conform to the norms are an indication that the rules of science governing the data are incomplete, and something new needs to be present to explain these unexpected outliers. The challenge of finding anomalies can be confounding since it requires codifying a complete knowledge of the known scientific behaviors and then projecting these known behaviors on the data to look for deviations. When utilizing machine learning, this presents a particular challenge since we require that the model not only understands scientific data perfectly but also recognizes when the data is inconsistent and out of the scope of its trained behavior. In this paper, we present three datasets aimed at developing machine learning-based anomaly detection for disparate scientific domains covering astrophysics, genomics, and polar science. We present the different datasets along with a scheme to make machine learning challenges around the three datasets findable, accessible, interoperable, and reusable (FAIR). Furthermore, we present an approach that generalizes to future machine learning challenges, enabling the possibility of large, more compute-intensive challenges that can ultimately lead to scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。