TLM模型宣称的泛化能力实为评估漏洞所致,真实表现远低于宣称。
The Illusion of Generalization in Tabular Language Models
- 用165个数据集重测Tabula-8B,发现其泛化能力主要依赖污染数据
- 二分类任务性能仅略高于多数类基线,四分位分类主导整体表现
- 指令微调无需表格训练即可恢复92.2%性能,提示泛化或源于格式熟悉度
Tabular Language Models(TLMs)被宣称具备强大的表格预测泛化能力。我们以代表性模型Tabula-8B为例,基于UniPredict基准中的165个数据集进行系统性再评估。研究发现:第一,二分类与类别分类任务的中位提升接近零,整体强性能完全由四分位分类任务驱动;第二,表现最优的数据集存在广泛污染,包括完整的训练测试重叠及规避标准去重机制的任务级泄漏;第三,未经表格数据暴露的指令微调可恢复92.2%的标准分类性能,而对四分位分类,格式熟悉度弥补了71.3%的性能差距,剩余差异归因于污染数据。这些结果表明,声称的泛化能力可能源于评估漏洞而非真正的表格推理能力。论文最后提出强化TLM评估的建议。
原文摘要 · Abstract (English)
Tabular Language Models (TLMs) have been claimed to achieve strong generalization for tabular prediction. We conduct a systematic re-evaluation of Tabula-8B as a representative TLM, utilizing 165 datasets from the UniPredict benchmark. Our investigation reveals three findings. First, binary and categorical classification achieve near-zero median lift over majority-class baselines and strong aggregate performance is driven entirely by quartile classification tasks. Second, top-performing datasets exhibit pervasive contamination, including complete train-test overlap and task-level leakage that evades standard deduplication. Third, instruction-tuning without tabular exposure recovers 92.2% of standard classification performance and on quartile classification, format familiarity closes 71.3% of the gap with the residual attributable to contaminated datasets. These findings suggest claimed generalization likely reflects evaluation artifacts rather than learned tabular reasoning. We conclude with recommendations for strengthening TLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。