arXiv:2603.04321cs.CVcs.AI2026-03

针对表格数据的少样本增量学习,提出新框架SPRINT提升性能并防止遗忘。

SPRINT: Semi-supervised Prototypical Representation for Few-Shot Class-Incremental Tabular Learning

  • 用置信度伪标签融合无标注数据,增强新类别表征
  • 在6个领域基准上达77.37%准确率(5样本),领先基线4.45%
  • 适合需要持续学习、标注少且存储成本低的场景

现实系统需在数据有限的情况下持续适应新概念,同时不遗忘已有知识。尽管少样本增量学习(FSCIL)在计算机视觉中已成熟,但其在表格数据领域的应用仍基本空白。与图像不同,表格数据流(如日志、传感器)具有大量未标注数据、专家标注稀缺和存储成本极低的特点,而现有基于视觉的方法因依赖固定缓冲区而忽略这些优势。本文提出SPRINT,首个专为表格分布设计的FSCIL框架。SPRINT采用混合历元训练策略,利用置信度伪标签丰富新类表征,并借助低存储成本保留基础类历史信息。在涵盖网络安全、医疗健康和生态等六个不同领域的基准上进行广泛评估,验证了SPRINT跨领域鲁棒性。其平均准确率达77.37%(5样本),优于最强增量基线4.45%。

原文摘要 · Abstract (English)

Real-world systems must continuously adapt to novel concepts from limited data without forgetting previously acquired knowledge. While Few-Shot Class-Incremental Learning (FSCIL) is established in computer vision, its application to tabular domains remains largely unexplored. Unlike images, tabular streams (e.g., logs, sensors) offer abundant unlabeled data, a scarcity of expert annotations and negligible storage costs, features ignored by existing vision-based methods that rely on restrictive buffers. We introduce SPRINT, the first FSCIL framework tailored for tabular distributions. SPRINT introduces a mixed episodic training strategy that leverages confidence-based pseudo-labeling to enrich novel class representations and exploits low storage costs to retain base class history. Extensive evaluation across six diverse benchmarks spanning cybersecurity, healthcare, and ecological domains, demonstrates SPRINT's cross-domain robustness. It achieves a state-of-the-art average accuracy of 77.37% (5-shot), outperforming the strongest incremental baseline by 4.45%.

少样本学习增量学习表格数据半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。