arXiv:2605.08519cs.LG2026-05

不依赖数据增强的表格数据半监督少样本学习新方法

SeBA: Semi-supervised few-shot learning via Separated-at-Birth Alignment for tabular data

  • 将表格数据分为两组独立视图,通过近邻对应关系对齐表征
  • 在多个基准数据集上达到当前最佳性能,提升特征与标签关联性
  • 适合医疗、金融等标签稀缺的表格数据分析场景

从少量标注数据和大量未标注样本中学习,即半监督少样本学习(SS-FSL),在医学、金融、科学等领域的表格数据应用中仍至关重要。现有方法多依赖为视觉或语言设计的自监督学习(SSL)框架,假设存在自然的数据增强方式。但表格数据定义有意义的增强手段非常困难,且易破坏语义,限制了传统SSL效果。本文重新思考表格数据的自监督学习,提出分离出生对齐(SeBA)方法,一种无需依赖增强的联合嵌入框架。核心思想是将数据分为两个独立但互补的视图,将一个视图的表示对齐至另一视图的最近邻对应关系。实验评估结合理论分析表明,SeBA生成的输出空间能有效改善特征-标签关系。在多个基准数据集上的实验证明,其多数情况下达到当前最优性能,为表格数据的半监督少样本学习开辟了新路径。

原文摘要 · Abstract (English)

Learning from scarce labeled data with a larger pool of unlabeled samples, known as semi-supervised few-shot learning (SS-FSL), remains critical for applications involving tabular data in domains like medicine, finance, and science. The existing SS-FSL methods often rely on self-supervised learning (SSL) frameworks developed for vision or language, which assume the availability of a natural form of data augmentations. For tabular data, defining meaningful augmentations is non-trivial and can easily distort semantics, limiting the effectiveness of conventional SSL. In this work, we rethink SSL for tabular data and propose Separated-at-Birth Alignment (SeBA), a joint-embedding framework for SS-FSL that eliminates the dependence on augmentations. Our core idea is to separate the data into two independent, but complementary views and align the representations of one view to mirror the nearest-neighbor correspondence of the data in the second view. Our experimental evaluation supported by a theoretical analysis justifies that SeBA generates an output space, which improves the feature-label relationship. An experimental study conducted in various benchmark datasets demonstrates that SeBA achieves the state-of-the-art performance in the majority of cases, opening a new avenue for SS-FSL paradigm in the domain of tabular data.

少样本学习半监督表格数据自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。