arXiv:2510.17385cs.LGcs.AI2025-10被引 1

用结构先验增强LLM,让其在表格预测上超越专业模型。

Strengthening LLMs for Tabular Prediction with Structural Priors

  • 通过列置换不变性设计强化学习训练,提升优化信号密度
  • 8B模型在139个数据集上达到强基线水平,零样本性能领先
  • 显著优于更大通用模型,最高提升53.17%(对比DeepSeek-R1)

表格预测长期由梯度提升决策树和专用深度表格模型主导,尽管大语言模型(LLMs)具备跨任务适应性和可解释推理能力,但难以实现竞争力。本文通过在LLM后训练中引入表格结构先验来填补这一差距。提出排列相对策略优化(PRPO),通过保持标签的列置换操作与两级优势估计,实现列置换不变性。该设计将稀疏的结果奖励转化为更密集且稳定的优化信号。在139个OpenML数据集上的大量实验表明,我们的8B模型达到了真正具有竞争力的性能水平,超越了强大的专用表格基线。它在全监督设置下表现强劲,在零样本场景中占据主导地位,并与32样本强基线表现相当。此外,其性能显著优于更大规模的通用及推理类LLM,相较DeepSeek-R1(685B)最高提升达53.17%。结果表明,基于结构先验的强化学习后训练是使LLM在表格预测中具备竞争力的有效路径。

原文摘要 · Abstract (English)

Tabular prediction has long been dominated by gradient-boosted decision trees and specialized deep tabular models, while large language models (LLMs) remain difficult to make competitive despite their cross-task adaptability and transparent reasoning traces. We address this gap by incorporating tabular structural priors into LLM post-training. Specifically, we propose Permutation Relative Policy Optimization (PRPO), which operationalizes column-permutation invariance through label-preserving column permutations and two-level advantage estimation. This design converts sparse outcome rewards into denser and more stable optimization signals. Extensive experiments on 139 OpenML datasets show that our 8B model reaches a genuinely competitive regime against strong specialized tabular baselines. It achieves strong fully supervised performance, dominates zero-shot settings, and performs on par with 32-shot strong baselines. Moreover, it substantially outperforms much larger general-purpose and reasoning LLMs, including up to a 53.17% improvement over DeepSeek-R1 (685B). These results show that structural-prior RL post-training is an effective route for making LLMs competitive in tabular prediction.

表格预测强化学习大模型结构先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。