arXiv:2502.11262cs.DBcs.AI2025-02被引 1

用多目标优化发现更均衡的高质量数据集,避免单一标准带来的偏差。

Generating Skyline Datasets for Data Science Models

  • 基于用户自定义性能指标,从多个数据源中构建综合表现优的数据集。
  • 提出三种算法,可高效生成在多个指标上均表现良好的数据集。
  • 适合需要多维度评估数据质量的数据科学流程优化场景。

为各类数据驱动的人工智能与机器学习模型准备高质量数据集已成为数据分析的核心任务。传统数据发现方法通常依赖单一预设的质量度量,可能导致下游任务出现偏差。本文提出MODis框架,通过优化多个用户自定义的模型性能指标来发现数据集。给定一组数据源和一个模型,MODis选择并整合数据源形成“天际线数据集”,使模型在所有性能指标上均达到预期表现。我们将MODis建模为多目标有限状态转换器,并推导出三种可行算法:第一种采用“从全集减少”策略,从通用模式开始逐步剔除低潜力数据;第二种引入双向策略,交替进行数据增强与缩减以降低计算成本;还提出了去偏算法以缓解天际线数据集中的偏差问题。实验验证了算法在效率与有效性上的优势,并展示了其在优化数据科学流水线中的应用价值。

原文摘要 · Abstract (English)

Preparing high-quality datasets required by various data-driven AI and machine learning models has become a cornerstone task in data-driven analysis. Conventional data discovery methods typically integrate datasets towards a single pre-defined quality measure that may lead to bias for downstream tasks. This paper introduces MODis, a framework that discovers datasets by optimizing multiple user-defined, model-performance measures. Given a set of data sources and a model, MODis selects and integrates data sources into a skyline dataset, over which the model is expected to have the desired performance in all the performance measures. We formulate MODis as a multi-goal finite state transducer, and derive three feasible algorithms to generate skyline datasets. Our first algorithm adopts a "reduce-from-universal" strategy, that starts with a universal schema and iteratively prunes unpromising data. Our second algorithm further reduces the cost with a bi-directional strategy that interleaves data augmentation and reduction. We also introduce a diversification algorithm to mitigate the bias in skyline datasets. We experimentally verify the efficiency and effectiveness of our skyline data discovery algorithms, and showcase their applications in optimizing data science pipelines.

数据发现多目标优化数据集构建机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。