用512个原型压缩数据,让表格大模型在无GPU下仍高效准确。
Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data

- 自适应选择关键样本,兼顾类别平衡与特征多样性。
- 仅用512个原型即保留原始性能,压缩比达1953倍。
- 无需重训练,在普通CPU上就能处理百万级表格数据。
预训练表格基础模型展现出强大的预测能力,但其在大规模数据集上的应用受限于有限的推理上下文。本文提出平衡自适应原型选择(BAPS),一种构建紧凑、信息保留上下文的框架,以实现可扩展的TabPFN推理。BAPS在不修改或重新训练预训练模型的前提下,联合保持代表性结构、有信息量的决策边界、局部密度、类别平衡和特征空间多样性。在包含百万行的HIGGS和SUSY数据集上的实验表明,使用512个原型即可维持强预测性能与可靠校准,对应约1953倍的上下文压缩。所有实验均在配备16GB内存的Intel Core i7 CPU上完成,未使用GPU加速。这些发现确立了有效的上下文构建为扩展预训练表格基础模型至百万规模数据集的实用机制。
原文摘要 · Abstract (English)
Pretrained tabular foundation models have demonstrated strong predictive capability; however, their application to large-scale datasets remains constrained by the limited inference context. This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference. Without modifying or retraining the pretrained model, BAPS jointly preserves representative structure, informative decision boundaries, local density, class balance, and feature-space diversity. Experiments on the million-row HIGGS and SUSY datasets show that 512 prototypes retain strong predictive performance and reliable calibration, corresponding to an approximately 1,953-fold context compression. All experiments were conducted on an Intel Core i7 CPU with 16 GB RAM and no GPU acceleration. These findings establish effective context construction as a practical mechanism for extending pretrained tabular foundation models to million-scale datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。