arXiv:2502.10125cs.LGcs.AI2025-02

无需共享特征或对齐数据,实现跨表格关系学习

Learning Relational Tabular Data without Shared Features

  • 通过软对齐机制利用损失差异判断数据匹配度
  • 在5个真实与5个合成数据集上提升性能最高达26.8%
  • 适合无共享特征的跨表学习场景,尤其适用于大规模表格

关系型表格数据的学习近年来受到广泛关注,但多数研究集中于单张表格,忽视了跨表学习的潜力。在缺乏共享特征和预对齐数据的场景下,跨表学习虽有巨大前景,却面临对齐空间庞大、准确对齐困难等挑战。我们提出隐式实体对齐学习(Leal)框架,可在无需共享特征或预对齐数据的情况下实现有效的跨表训练。Leal基于正确对齐的数据比错误对齐数据产生更低损失的原理,采用软对齐机制,并结合可微分聚类采样模块,确保在大规模关系表格上的高效扩展。此外,我们提供了该模块近似能力的理论证明。在五个真实数据集和五个合成数据集上的大量实验表明,Leal相比当前最优方法预测性能最高提升26.8%,验证了其有效性和可扩展性。

原文摘要 · Abstract (English)

Learning relational tabular data has gained significant attention recently, but most studies focus on single tables, overlooking the potential of cross-table learning. Cross-table learning, especially in scenarios where tables lack shared features and pre-aligned data, offers vast opportunities but also introduces substantial challenges. The alignment space is immense, and determining accurate alignments between tables is highly complex. We propose Latent Entity Alignment Learning (Leal), a novel framework enabling effective cross-table training without requiring shared features or pre-aligned data. Leal operates on the principle that properly aligned data yield lower loss than misaligned data, a concept embodied in its soft alignment mechanism. This mechanism is coupled with a differentiable cluster sampler module, ensuring efficient scaling to large relational tables. Furthermore, we provide a theoretical proof of the cluster sampler's approximation capacity. Extensive experiments on five real-world and five synthetic datasets show that Leal achieves up to a 26.8% improvement in predictive performance compared to state-of-the-art methods, demonstrating its effectiveness and scalability.

跨表学习表格数据无监督对齐关系建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。