无需参数调优的编码器仍可高效预测数据库中的缺失值
Parameter-Free Encoders Remain Viable for RDB Foundation Models

- 使用无需训练的简单编码器处理关系型数据库
- 在多个基准任务中表现接近顶尖水平
- 适合快速部署于企业级数据预测场景
面对存储异构表格数据的关系型数据库(RDB),如何预测目标列中的缺失(或未来)数值?由于企业环境中目标列种类繁多,每次新任务都从头训练模型不现实。基于RDB专用编码器的冻结基础模型提供可行方案,但理想设计仍不明朗。一方面,近期研究指出某些无参数子图编码器结合单表基础模型,无需RDB特有预训练即可达到接近最先进性能;另一方面,也有研究主张使用可训练编码器并利用可观测标签进行任务特定表示学习。为解决这一分歧,本文分析了在标签作为输入时RDB编码器的特性,证明了可训练参数的潜在局限性。实证上,我们展示更简单的无参数编码器在多个相关基准任务中仍能保持强性能。
原文摘要 · Abstract (English)
Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of interest? As the space of potential targets is vast across enterprise settings, it is preferable to avoid learning a new model from scratch each time there is a new prediction task. Frozen foundation models based on RDB-specific encoders provide a viable solution, but ideal design remains an open question. On the one hand, it has recently been argued that certain parameter-free subgraph encoders combined with single-table foundation models can achieve near SOTA performance, with no RDB-specific pre-training required. Meanwhile, other contemporary studies advocate for parameterized encoders pre-trained to exploit observable labels for learning task-specific representations. To address this ambiguity, we analyze RDB encoder properties specifically when labels are present as inputs, proving limitations on the potential efficacy of trainable encoder parameters. As empirical validation, we demonstrate that considerably simpler parameter-free encoders are still capable of strong performance across many relevant benchmarking tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。