arXiv:2409.10139cs.DBcs.AI2024-09

无需领域知识,自动检测并修复数据缺失、重复和不一致问题。

Towards Explainable Automated Data Quality Enhancement without Domain Knowledge

  • 结合统计方法与机器学习,实现可解释的数据质量评估。
  • 有效识别并修正缺失值、重复项和拼写错误。
  • 适合需要透明可信数据清洗的科研与工业场景。

在大数据时代,确保数据集质量在各个领域变得日益重要。本文提出一个全面框架,可自动评估并修正任意数据集中的数据质量问题,无论其具体内容如何,重点关注文本和数值数据。主要目标是解决三类基本缺陷:缺失、冗余和不一致。核心在于强调可解释性和可理解性,确保数据异常识别与修正的逻辑透明可懂。采用混合方法,融合统计技术与机器学习算法,在准确率与可解释性之间取得平衡,使用户能够信任并理解评估过程。针对自动化数据质量评估在时效性和准确性方面的挑战,采取务实策略:仅在必要时使用资源密集型算法,优先选择简单高效的解决方案。通过在公开数据集上的实际分析,展示了在保持可解释性条件下提升数据质量所面临的挑战。实证表明,该方法在检测并修正缺失值、重复项和拼写错误方面有效,但在统计离群值和逻辑错误的处理上,仍需进一步突破现有约束条件下的准确率瓶颈。

原文摘要 · Abstract (English)

In the era of big data, ensuring the quality of datasets has become increasingly crucial across various domains. We propose a comprehensive framework designed to automatically assess and rectify data quality issues in any given dataset, regardless of its specific content, focusing on both textual and numerical data. Our primary objective is to address three fundamental types of defects: absence, redundancy, and incoherence. At the heart of our approach lies a rigorous demand for both explainability and interpretability, ensuring that the rationale behind the identification and correction of data anomalies is transparent and understandable. To achieve this, we adopt a hybrid approach that integrates statistical methods with machine learning algorithms. Indeed, by leveraging statistical techniques alongside machine learning, we strike a balance between accuracy and explainability, enabling users to trust and comprehend the assessment process. Acknowledging the challenges associated with automating the data quality assessment process, particularly in terms of time efficiency and accuracy, we adopt a pragmatic strategy, employing resource-intensive algorithms only when necessary, while favoring simpler, more efficient solutions whenever possible. Through a practical analysis conducted on a publicly provided dataset, we illustrate the challenges that arise when trying to enhance data quality while keeping explainability. We demonstrate the effectiveness of our approach in detecting and rectifying missing values, duplicates and typographical errors as well as the challenges remaining to be addressed to achieve similar accuracy on statistical outliers and logic errors under the constraints set in our work.

数据清洗可解释性自动化数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。