KumoRFM-2可直接处理多表关系数据,少样本下超越传统方法8%。
KumoRFM-2: Scaling Foundation Models for Relational Learning
- 直接建模多表关系,无需人工展平数据
- 在41个基准上比监督模型最高提升8%,冷启动和噪声下仍稳定
- 支持少样本学习与微调,适用于超大规模关系数据
我们提出KumoRFM-2,下一代用于关系数据的预训练基础模型。该模型支持上下文学习与微调,适用于多种预测任务。与表格类基础模型不同,KumoRFM-2原生处理关系数据,可同时处理一个或多个连接表,无需手动展平或生成目标变量,且保持时间一致性。其通过合成与真实世界数据在四个维度进行预训练:单表的行与列维度,以及数据库级别的外键与跨样本维度。相比前代,KumoRFM-2更早注入任务信息,提升对相关列的选择精度并增强对噪声数据的鲁棒性。在41个挑战性基准上的实验表明,其性能比监督与基础方法最高提升8%,并在极端冷启动与噪声环境下保持强表现。据我们所知,这是首个在常见基准任务上少样本即超越监督方法的基础模型,微调后性能进一步提升。此前的KumoRFM-1仅限于小规模内存数据集,而KumoRFM-2已扩展至百亿级关系数据集。
原文摘要 · Abstract (English)
We introduce KumoRFM-2, the next iteration of a pre-trained foundation model for relational data. KumoRFM-2 supports in-context learning as well as fine-tuning and is applicable to a wide range of predictive tasks. In contrast to tabular foundation models, KumoRFM-2 natively operates on relational data, processing one or more connected tables simultaneously without manual table flattening or target variable generation, all while preserving temporal consistency. KumoRFM-2 leverages a large corpus of synthetic and real-world data to pre-train across four axes: the row and column dimensions at the individual table level, and the foreign key and cross-sample dimensions at the database level. In contrast to its predecessor, KumoRFM-2 injects task information as early as possible, enabling sharper selection of task-relevant columns and improved robustness to noisy data. Through extensive experiments on 41 challenging benchmarks and analysis around expressivity and sensitivity, we demonstrate that KumoRFM-2 outperforms supervised and foundational approaches by up to 8%, while maintaining strong performance under extreme settings of cold start and noisy data. To our knowledge, this is the first time a few-shot foundation model has been shown to surpass supervised approaches on common benchmark tasks, with performance further improving upon fine-tuning. Finally, while KumoRFM-1 was limited to small-scale in-memory datasets, KumoRFM-2 scales to billion-scale relational datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。