构建可生长的生物医学表格数据集基准,评测不同模型在高维小样本下的表现。
TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables
- 建立43个跨领域的生物医学表格数据集,统一交叉验证协议进行对比
- 1万特征100样本时,RealTabPFN v2.5表现最优,领先逻辑回归145 Elo
- 支持结果可复现、可提交新数据,适合医疗表结构学习研究者使用
生物医学表格常包含数千个变量却仅有数十或数百个标注样本,这一场景在通用表格基准中缺乏代表。我们提出TabBench-Bio,一个动态交互式基准,包含43个跨领域的生物医学数据集。在统一交叉验证协议下,我们在28个特征-样本组合上比较了经典估计器、神经网络与表格基础模型。在参考场景(10,000特征,100训练样本)中,RealTabPFN v2.5点估计最高,次为逻辑回归与TabDPT,二者点估计几乎相同。配对自助法显示,RealTabPFN v2.5相较逻辑回归领先145 Elo(95%置信区间[59, 232])。表格基础模型整体表现领先,但最优配置依赖于操作点与生物医学模态。采用一小时“极端”预设的AutoGluon作为资源密集型参照,报告其在参考点的表现。折痕级预测、运行状态与确定性聚合确保所有结果可复现、可重用。我们邀请社区贡献:TabBench-Bio持续生长,欢迎提交未充分代表的检测方法与临床终点的新数据集,以纳入未来版本。交互式排行榜见:https://tabbench-bio.eu
原文摘要 · Abstract (English)
Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN v2.5 has the highest point estimate, followed by Logistic Regression and TabDPT, whose point estimates are nearly identical. A paired bootstrap over the target pool separates RealTabPFN v2.5 from Logistic Regression by 145 Elo (95% interval [59, 232]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour "extreme" preset, is configured as a separate resource-intensive reference and is reported here at the reference cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. We invite the community to contribute: TabBench-Bio is designed to grow, and we welcome submissions of new biomedical tabular datasets, particularly from underrepresented assays and clinical endpoints, for inclusion in future releases. The interactive leaderboard is available at: https://tabbench-bio.eu
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。