arXiv:2507.11960cs.HCcs.LG2025-07

用可视化工具辅助提升数据质量,让模型效果更好

d-DQIVAR: Data-centric Visual Analytics and Reasoning for Data Quality Improvement

  • 融合数据与流程双驱动,自动处理缺失、异常、重复等问题
  • 通过柯尔莫哥洛夫-斯米尔诺夫检验评估数据分布变化,确保优化有效
  • 适合数据科学家和领域专家在真实场景中改进数据质量

提升数据质量(DQ)的方法主要分为数据驱动和流程驱动两类。但以往研究多采用批量数据预处理,难以有效提升机器学习(ML)模型性能,且常导致数据特征失真。现有工作聚焦于数据预处理,而非真正的数据质量改进(DQI)。本文提出 d-DQIVAR,一个新型可视化分析系统,旨在支持切实的数据质量改进策略以提升模型表现。系统结合数据驱动与流程驱动方法:数据驱动部分处理缺失值填补、异常值检测、去重、格式标准化及特征选择;流程驱动部分基于数据质量维度与模型性能评估改进过程,并应用柯尔莫哥洛夫-斯米尔诺夫检验。通过案例研究、评估与用户实验,验证了系统能有效整合专家与领域知识,实现可操作的工作流。

原文摘要 · Abstract (English)

Approaches to enhancing data quality (DQ) are classified into two main categories: data- and process-driven. However, prior research has predominantly utilized batch data preprocessing within the data-driven framework, which often proves insufficient for optimizing machine learning (ML) model performance and frequently leads to distortions in data characteristics. Existing studies have primarily focused on data preprocessing rather than genuine data quality improvement (DQI). In this paper, we introduce d-DQIVAR, a novel visual analytics system designed to facilitate DQI strategies aimed at improving ML model performance. Our system integrates visual analytics techniques that leverage both data-driven and process-driven approaches. Data-driven techniques tackle DQ issues such as imputation, outlier detection, deletion, format standardization, removal of duplicate records, and feature selection. Process-driven strategies encompass evaluating DQ and DQI procedures by considering DQ dimensions and ML model performance and applying the Kolmogorov-Smirnov test. We illustrate how our system empowers users to harness expert and domain knowledge effectively within a practical workflow through case studies, evaluations, and user studies.

数据质量可视化分析机器学习数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。