提出新方法让表格数据学得更清晰,性能超越主流模型。
TabDeco: A Comprehensive Contrastive Framework for Decoupled Representations in Tabular Data
- 用注意力机制从行和列双角度解耦特征表示
- 在多个基准任务上优于XGBoost等经典算法
- 适合需要高质量表格表征的工业级建模场景
表示学习是现代人工智能的核心,推动了众多应用的进展。尽管自监督对比学习在计算机视觉和自然语言处理中取得显著突破,但其在表格数据中的应用面临独特挑战。传统方法多关注模型架构与损失函数优化,却忽视了从特征交互、实例模式和批次上下文等多个维度构建有意义的正负样本对。为此,我们提出TabDeco,一种基于注意力编码策略的新型方法,通过对比学习框架在特征、实例和数据批次三个层次上有效解耦表示。借助创新的特征解耦层级结构,TabDeco在多个基准任务中持续超越现有深度学习方法及主流梯度提升算法(如XGBoost、CatBoost、LightGBM),充分证明其在推进表格数据表示学习方面的有效性。
原文摘要 · Abstract (English)
Representation learning is a fundamental aspect of modern artificial intelligence, driving substantial improvements across diverse applications. While selfsupervised contrastive learning has led to significant advancements in fields like computer vision and natural language processing, its adaptation to tabular data presents unique challenges. Traditional approaches often prioritize optimizing model architecture and loss functions but may overlook the crucial task of constructing meaningful positive and negative sample pairs from various perspectives like feature interactions, instance-level patterns and batch-specific contexts. To address these challenges, we introduce TabDeco, a novel method that leverages attention-based encoding strategies across both rows and columns and employs contrastive learning framework to effectively disentangle feature representations at multiple levels, including features, instances and data batches. With the innovative feature decoupling hierarchies, TabDeco consistently surpasses existing deep learning methods and leading gradient boosting algorithms, including XG-Boost, CatBoost, and LightGBM, across various benchmark tasks, underscoring its effectiveness in advancing tabular data representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。