通过数据浓缩提升表格模型迁移能力,缓解数据分布差异问题。
Context-Constrained Transfer Learning for Tabular Foundation Models via Data Distillation
- 基于目标数据构建紧凑源上下文,优化覆盖与后验兼容性
- 在有限上下文下实现比直接拼接更优的迁移性能
- 适合处理源-目标数据分布不一致的表格迁移任务
表格基础模型(TFMs)在上下文学习中表现出色,但迁移学习受限于严格的上下文大小约束和源-目标任务间分布偏移。直接合并异构源数据可能导致负迁移。为此,我们提出基于锚定与蒸馏的上下文受限迁移学习框架(TL-ANDI)。该框架通过求解一个预算约束的最优传输问题,构建紧凑源上下文,其代价同时衡量目标协变量覆盖率与后验兼容性。选定的锚点样本经局部标签蒸馏,并结合目标数据进行残差校准,显著提升迁移效果。
原文摘要 · Abstract (English)
Tabular Foundation Models (TFMs) have demonstrated strong empirical performance as black-box inference engines through in-context learning. However, their use in transfer learning is limited by two obstacles: strict context-size constraints and sensitivity to distribution shifts between source and target tasks. Directly pooling heterogeneous source data can therefore lead to negative transfer. To address these challenges, we propose Context-Constrained Transfer Learning via ANchoring and DIstillation (TL-ANDI), a posterior-aware distillation framework for TFMs. TL-ANDI constructs a compact source context by solving a budget-constrained optimal transport problem whose cost jointly measures target covariate coverage and posterior compatibility. The selected anchor samples are then equipped with locally distilled labels and combined with a residual calibration step using target data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。