arXiv:2507.16833cs.LGphysics.data-an2025-07

提出一套自动检测修复自驱动实验室数据噪声的方法,提升材料发现精度。

Exploring the Limitations of kNN Noisy Feature Detection and Recovery for Self-Driving Labs

  • 用kNN方法自动识别数据中的异常特征并定位可修正样本
  • 高噪声强度和大训练集更易检测修复,连续分布特征恢复效果更好
  • 适用于小样本、噪声多的材料实验数据,为数据质量提供基准

自驱动实验室(SDL)通过结合机器学习与自动化实验平台,有望加速材料发现。然而,输入参数采集错误可能污染用于建模系统性能的特征,影响当前及后续实验。本研究开发了一套自动化工作流程,系统性地检测噪声特征、判断可修正的样本-特征配对,并恢复正确特征值。通过在密度泛函理论(DFT)和SDL数据集上系统分析数据集规模、噪声强度、噪声类型及特征值分布对噪声可检测性和可恢复性的影响发现:高噪声强度和大规模训练数据有助于噪声检测与修正;低强度噪声虽降低检测与恢复能力,但可通过更大规模的干净数据集补偿;连续且分散的特征分布比离散或窄分布的特征更具可恢复性。该研究不仅展示了在噪声、数据有限及特征分布差异下实现理性数据恢复的模型无关框架,还为材料数据集中kNN插补提供了可量化的基准。最终目标是提升自动化材料发现中的数据质量与实验精度。

原文摘要 · Abstract (English)

Self-driving laboratories (SDLs) have shown promise to accelerate materials discovery by integrating machine learning with automated experimental platforms. However, errors in the capture of input parameters may corrupt the features used to model system performance, compromising current and future campaigns. This study develops an automated workflow to systematically detect noisy features, determine sample-feature pairings that can be corrected, and finally recover the correct feature values. A systematic study is then performed to examine how dataset size, noise intensity, noise type, and feature value distribution affect both the detectability and recoverability of noisy features on both Density Functional Theory (DFT) and SDL datasets. In general, high-intensity noise and large training datasets are conducive to the detection and correction of noisy features. Low-intensity noise reduces detection and recovery but can be compensated for by larger clean training data sets. Detection and correction results vary between features, with continuous and dispersed feature distributions showing greater recoverability compared to features with discrete or narrow distributions. This systematic study not only demonstrates a model agnostic framework for rational data recovery in the presence of noise, limited data, and differing feature distributions but also provides a tangible benchmark of kNN imputation in materials datasets. Ultimately, it aims to enhance data quality and experimental precision in automated materials discovery.

数据修复材料发现噪声检测kNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。