arXiv:2605.28554cs.LG2026-05中稿 · ESANN 2026

TFM模型预测准但不确定度不准,可靠应用仍存挑战

High Performance, Low Reliability: Uncertainty Benchmarking for Tabular Foundation Models

论文配图:High Performance, Low Reliability: Uncertainty Benchmarking for Tabular Foundation Models
图 1 · 摘自论文原文
  • 用TALENT基准对比112个数据集上的TFM、GBDT等模型
  • TFM AUC最高但条件覆盖率比GBDT低,体现性能-可靠性权衡
  • 适合关注模型可信度的工业落地与风险敏感场景研究者

近期表格基础模型(TFMs)在预测性能上已超越梯度提升决策树(GBDTs),但其不确定性量化能力尚未得到充分重视。本研究通过在112个数据集组成的TALENT基准上,系统比较了TFMs、GBDTs与经典基线模型的表现。结果表明:尽管TFMs在AUC指标上达到最优,但在共形预测下的条件覆盖率(SSCS)却低于GBDTs,揭示出性能与不确定性之间的权衡。对合成数据集的补充实验进一步刻画了该现象加剧的条件。结论是,虽然TFMs推动了预测性能边界,但实现良好校准的不确定性仍是其可靠部署的重大挑战。代码已开源。

原文摘要 · Abstract (English)

Recent Tabular Foundation Models (TFMs) have demonstrated state-of-the-art predictive performance, often surpassing Gradient-Boosted Decision Trees (GBDTs). However, the trustworthiness of these models, particularly their uncertainty quantification, has been largely overlooked. We investigate this gap through an extensive study comparing TFMs, GBDTs, and classical baselines on the 112 datasets of the TALENT benchmark. Our results reveal a performance-uncertainty trade-off: although TFMs achieve the highest predictive performance, measured by AUC, they exhibit lower conditional coverage under conformal prediction, measured by SSCS, compared to GBDTs. Complementary experiments on synthetic datasets further characterize the regimes in which this effect intensifies. We conclude that while TFMs advance predictive frontiers, achieving well-calibrated uncertainty remains a major open challenge for their reliable adoption. Code is available at: https://github.com/jose-melo/high-performance-low-reliability

表格模型不确定性可信赖AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。