arXiv:2412.12864cs.LG2024-12AAAI被引 2

针对混合变量表格数据,提出基于测地流的半监督学习新方法。

Geodesic Flow Kernels for Semi-Supervised Learning on Mixed-Variable Tabular Dataset

  • 按变量类型设计特定扰动方式,更好模拟真实数据变化。
  • 用测地流核度量扰动前后数据几何差异,提升相似性判断。
  • 结合树结构嵌入,利用标签数据中的层级关系,适合小样本场景。

表格数据因同时包含连续与分类变量而具有异质性,现有方法难以有效捕捉其内在结构。本文提出 GFTab(Geodesic Flow Kernels for Semi-Supervised Learning on Mixed-Variable Tabular Dataset),一种专为混合变量表格数据设计的半监督学习框架。GFTab 包含三项关键创新:1)针对连续与分类变量特性设计的变量特异性扰动方法;2)基于测地流核的相似性度量,用于捕捉扰动输入间的几何变化;3)基于树的嵌入方法,利用已有标签数据中的层级关系。为严谨评估,我们构建了涵盖多个领域、规模和变量构成的 21 个表格数据集。实验结果表明,GFTab 在多数数据集上优于现有机器学习/深度学习模型,尤其在标签数据有限的场景下表现更优。

原文摘要 · Abstract (English)

Tabular data poses unique challenges due to its heterogeneous nature, combining both continuous and categorical variables. Existing approaches often struggle to effectively capture the underlying structure and relationships within such data. We propose GFTab (Geodesic Flow Kernels for Semi- Supervised Learning on Mixed-Variable Tabular Dataset), a semi-supervised framework specifically designed for tabular datasets. GFTab incorporates three key innovations: 1) Variable-specific corruption methods tailored to the distinct properties of continuous and categorical variables, 2) A Geodesic flow kernel based similarity measure to capture geometric changes between corrupted inputs, and 3) Tree-based embedding to leverage hierarchical relationships from available labeled data. To rigorously evaluate GFTab, we curate a comprehensive set of 21 tabular datasets spanning various domains, sizes, and variable compositions. Our experimental results show that GFTab outperforms existing ML/DL models across many of these datasets, particularly in settings with limited labeled data.

半监督学习表格数据测地流小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。