arXiv:2501.13905cs.LG2025-01被引 3

提出表格数据蒸馏新框架,提升小样本数据质量

On Learning Representations for Tabular Data Distillation

  • 基于列嵌入学习表示,解决表格数据异构性难题
  • 在7种模型上使蒸馏数据质量提升0.5%-143%
  • 适合需要高效低耗建模的工业场景

数据蒸馏从大规模数据集中生成少量信息密集的实例,可降低存储、隐私及计算成本,但现有研究多聚焦图像数据。本文研究表格数据蒸馏,面临特征异构性以及决策树集成、近邻预测等不可微模型的挑战。为此,提出基于列嵌入表示学习的蒸馏框架 $ exttt{TDColER}$,并构建了表格数据蒸馏基准 ${f iny TDBench}$。基于该基准的详尽评估共生成226,890个蒸馏数据集,训练548,880个模型,结果表明 $ exttt{TDColER}$ 可在7种不同表格学习模型上将现有蒸馏方案的数据质量提升0.5%至143%。

原文摘要 · Abstract (English)

Dataset distillation generates a small set of information-rich instances from a large dataset, resulting in reduced storage requirements, privacy or copyright risks, and computational costs for downstream modeling, though much of the research has focused on the image data modality. We study tabular data distillation, which brings in novel challenges such as the inherent feature heterogeneity and the common use of non-differentiable learning models (such as decision tree ensembles and nearest-neighbor predictors). To mitigate these challenges, we present $\texttt{TDColER}$, a tabular data distillation framework via column embeddings-based representation learning. To evaluate this framework, we also present a tabular data distillation benchmark, ${\sf \small TDBench}$. Based on an elaborate evaluation on ${\sf \small TDBench}$, resulting in 226,890 distilled datasets and 548,880 models trained on them, we demonstrate that $\texttt{TDColER}$ is able to boost the distilled data quality of off-the-shelf distillation schemes by 0.5-143% across 7 different tabular learning models.

数据蒸馏表格数据表示学习机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。