arXiv:2501.14717cs.CL2025-01Conference of the …被引 2

对比12个表格大模型,发现选基模型比选数据更重要

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects

  • 用3个基模型+4个数据集复现12个表格LLM
  • 基模型差异导致性能差距远超数据影响
  • 适合研究表格建模与指令微调的开发者

表格建模已发展数十年。本文重新审视这一进程,在大模型时代揭示新挑战,尤其是指令微调中基模型与训练数据多样导致的性能归因难题。我们通过在四个现有数据集上对三个基础模型进行指令微调,复现了四个表格大模型,共获得12个模型,并在16个表格基准上进行评估。这是首个定量分离训练数据与基模型影响的研究,结果表明基模型选择的作用远大于训练数据本身。泛化与推理能力仍具挑战,亟需未来工作改进。基于此,我们提出表格建模的未来方向建议。

原文摘要 · Abstract (English)

Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly the paradox of choice: the difficulty of attributing performance gains amid diverse base models and training sets in the context of table instruction tuning. We replicate four table LLMs by instruction-tuning three foundation models on four existing datasets, yielding 12 models. We then evaluate these models across 16 table benchmarks. Our study is the first to quantitatively disentangle the effects of training data and base model selection, revealing that base model choice plays a more dominant role than the training data itself. Generalization and reasoning remain challenging, inviting future effort on table modeling. Based on our findings, we share our thoughts on the future directions for table modeling.

表格建模大模型指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。