用无监督方法自动识别数据质量,提升机器学习效果
Enhancing Machine Learning Performance through Intelligent Data Quality Assessment: An Unsupervised Data-centric Framework
- 结合质量度量与无监督学习,自动区分高低质量数据
- 在三个反义寡核苷酸数据集上提升模型性能
- 适合数据预处理耗时长的科研与工业场景
数据质量差会限制机器学习的优势,并削弱高性能机器学习系统的表现。随着数据量和复杂性的增加,数据质量问题愈发突出,导致机器学习流程前需投入大量时间进行数据准备与清洗。为此,本文提出一种智能数据驱动评估框架,通过整合质量度量与无监督学习,实现对高质与低质数据的区分,从而提升机器学习系统的性能。该框架具备灵活性与通用性,适用于多种领域与应用场景。我们在分析化学领域的实际案例中进行了验证,测试了三个反义寡核苷酸数据集。借助领域专家定义质量指标并评估结果,发现该质量导向的数据评估框架能有效识别高质量数据特征,指导高效实验设计,进而显著提升机器学习系统表现。
原文摘要 · Abstract (English)
Poor data quality limits the advantageous power of Machine Learning (ML) and weakens high-performing ML software systems. Nowadays, data are more prone to the risk of poor quality due to their increasing volume and complexity. Therefore, tedious and time-consuming work goes into data preparation and improvement before moving further in the ML pipeline. To address this challenge, we propose an intelligent data-centric evaluation framework that can identify high-quality data and improve the performance of an ML system. The proposed framework combines the curation of quality measurements and unsupervised learning to distinguish high- and low-quality data. The framework is designed to integrate flexible and general-purpose methods so that it is deployed in various domains and applications. To validate the outcomes of the designed framework, we implemented it in a real-world use case from the field of analytical chemistry, where it is tested on three datasets of anti-sense oligonucleotides. A domain expert is consulted to identify the relevant quality measurements and evaluate the outcomes of the framework. The results show that the quality-centric data evaluation framework identifies the characteristics of high-quality data that guide the conduct of efficient laboratory experiments and consequently improve the performance of the ML system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。