arXiv:2510.06162cs.LG2025-10被引 11

通过合成数据继续预训练,让模型轻松处理超3万维生物特征。

TabPFN-Wide: Continued Pre-Training for Extreme Feature Counts

  • 用定制先验生成合成数据,对基础模型进行持续预训练。
  • 在真实组学数据上,识别出与已有生物学发现重合的高相关特征。
  • 支持超大规模特征输入且保持可解释性,适合生物医学研究。

从分子测量与病理关系中挖掘新洞见是机器学习在生物医学中的重要应用。该领域数据通常观测样本极少但特征数达数千,传统表格机器学习方法难以应对。尽管已有预训练网络作为表格数据预测的基础模型,但当前无法处理超过500维的特征。虽可通过特征降维使用,但会损害特征重要性分析。本文提出一种策略:在自定义先验采样的合成数据上对现有模型进行持续预训练。所得模型TabPFN-Wide在性能上达到或超越基线模型,同时对噪声更具鲁棒性。它可无缝扩展至30,000+个类别与连续特征,无论噪声水平如何,仍保持固有可解释性,这对生物医学应用至关重要。实验证明,模型识别出的多数关键特征与先前生物学发现一致,部分则为未来研究提供潜在起点。

原文摘要 · Abstract (English)

Revealing novel insights from the relationship between molecular measurements and pathology remains a very impactful application of machine learning in biomedicine. Data in this domain typically contain only a few observations but thousands of potentially noisy features, posing challenges for conventional tabular machine learning approaches. While prior-data fitted networks emerge as foundation models for predictive tabular data tasks, they are currently not suited to handle large feature counts (>500). Although feature reduction enables their application, it hinders feature importance analysis. We propose a strategy that extends existing models through continued pre-training on synthetic data sampled from a customized prior. The resulting model, TabPFN-Wide, matches or exceeds its base model's performance, while exhibiting improved robustness to noise. It seamlessly scales beyond 30,000 categorical and continuous features, regardless of noise levels, while maintaining inherent interpretability, which is critical for biomedical applications. Our results demonstrate that prior-informed adaptation is suitable to enhance the capability of foundation models for high-dimensional data. On real-world omics datasets, we show that many of the most relevant features identified by the model overlap with previous biological findings, while others propose potential starting points for future studies.

生物医学表格模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。