arXiv:2501.09591cs.LG2025-01中稿 · 2025 SIAM Internat…被引 6

提出两种新度量方法,高效评估数据集间相似性。

Metrics for Inter-Dataset Similarity with Example Applications in Synthetic Data and Feature Selection Evaluation -- Extended Version

  • 基于全局数据分布设计双度量框架,避免参数敏感
  • 在合成数据与特征选择评估中验证有效性
  • 适合需要可靠数据相似性判断的研究者使用

在机器学习与数据挖掘中,衡量数据集间相似性是一项重要任务,具有多种应用场景。现有方法存在计算成本高、适用范围有限或对参数选择敏感等问题,且缺乏对数据集整体结构的全面考量。本文提出两种新的数据集间相似性度量方法,探讨其数学基础与理论依据。通过在合成数据评估与特征选择方法评价两个应用中验证其有效性,理论与实证研究均表明所提方法具有优越性能。

原文摘要 · Abstract (English)

Measuring inter-dataset similarity is an important task in machine learning and data mining with various use cases and applications. Existing methods for measuring inter-dataset similarity are computationally expensive, limited, or sensitive to different entities and non-trivial choices for parameters. They also lack a holistic perspective on the entire dataset. In this paper, we propose two novel metrics for measuring inter-dataset similarity. We discuss the mathematical foundation and the theoretical basis of our proposed metrics. We demonstrate the effectiveness of the proposed metrics by investigating two applications in the evaluation of synthetic data and in the evaluation of feature selection methods. The theoretical and empirical studies conducted in this paper illustrate the effectiveness of the proposed metrics.

数据相似性合成数据特征选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。