用对比学习自动发现科学数据中的异常,支持可量化的科学发现声明。
AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing

- 通过对比预训练生成低维数据表示,提升异常检测表达能力。
- 在多领域数据中成功检测到微小异常,统计显著性可量化。
- 适合需要严格统计验证的科研场景,如天体物理与生物实验。
大规模科学数据中的新颖性检测面临两大挑战:实验数据噪声大、维度高,且需对发现的异常做出统计上可靠的陈述。尽管已有大量基于降维的异常检测研究,但多数方法无法输出可量化科学发现结论。本文提出首个面向科学严谨性需求的新颖性检测统一流程——AutoSciDACT(自动化科学发现中的异常对比测试)。该方法首先利用对比学习,在多个科学领域丰富的高质量模拟数据和专家指导的数据增强策略基础上,构建具有表达力的低维数据嵌入。随后,这些紧凑嵌入被用于基于机器学习的两样本检验,采用新物理学习机(NPLM)框架,精准识别并统计量化观测数据相对于参考分布(零假设)的偏差。我们在天文、物理、生物、图像及合成数据集上进行实验,证明了该方法在所有领域均能对微小异常注入保持高度敏感。
原文摘要 · Abstract (English)
Novelty detection in large scientific datasets faces two key challenges: the noisy and high-dimensional nature of experimental data, and the necessity of making statistically robust statements about any observed outliers. While there is a wealth of literature on anomaly detection via dimensionality reduction, most methods do not produce outputs compatible with quantifiable claims of scientific discovery. In this work we directly address these challenges, presenting the first step towards a unified pipeline for novelty detection adapted for the rigorous statistical demands of science. We introduce AutoSciDACT (Automated Scientific Discovery with Anomalous Contrastive Testing), a general-purpose pipeline for detecting novelty in scientific data. AutoSciDACT begins by creating expressive low-dimensional data representations using a contrastive pre-training, leveraging the abundance of high-quality simulated data in many scientific domains alongside expertise that can guide principled data augmentation strategies. These compact embeddings then enable an extremely sensitive machine learning-based two-sample test using the New Physics Learning Machine (NPLM) framework, which identifies and statistically quantifies deviations in observed data relative to a reference distribution (null hypothesis). We perform experiments across a range of astronomical, physical, biological, image, and synthetic datasets, demonstrating strong sensitivity to small injections of anomalous data across all domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。