arXiv:2509.16889cs.CL2025-09EMNLP被引 13

用强化学习提升复杂表格理解能力,效果超越传统方法。

Can GRPO Boost Complex Multimodal Table Understanding?

  • 三阶段强化学习框架,逐步优化表格感知与推理。
  • 在多个数据集上显著超越监督微调和原有RL方法。
  • 适合需要强逻辑推理的多模态表格任务研究者。

现有表格理解方法受限于复杂结构和深层逻辑推理。尽管监督微调(SFT)占主导,但强化学习(如分组相对策略优化,GRPO)虽有潜力,却因初始策略准确率低和奖励粗糙而表现不佳。本文提出Table-R1,一种三阶段强化学习框架:(1) 预热阶段激发初始感知与推理能力;(2) 感知对齐GRPO(PA-GRPO),采用连续树编辑距离相似性(TEDS)奖励以识别表格结构与内容;(3) 提示补全GRPO(HC-GRPO),基于提示引导问题使用细粒度残差步数奖励。大量实验表明,Table-R1在保留与外部数据集上均显著提升模型表格推理性能,大幅优于SFT与原始GRPO。值得注意的是,Qwen2-VL-7B结合Table-R1超越更大规模专用表格模型(如Table-LLaVA 13B),甚至在保留数据集上接近闭源模型GPT-4o的表现,验证了各阶段对克服初始化瓶颈与奖励稀疏性的有效性,推动了鲁棒多模态表格理解的发展。

原文摘要 · Abstract (English)

Existing table understanding methods face challenges due to complex table structures and intricate logical reasoning. While supervised finetuning (SFT) dominates existing research, reinforcement learning (RL), such as Group Relative Policy Optimization (GRPO), has shown promise but struggled with low initial policy accuracy and coarse rewards in tabular contexts. In this paper, we introduce Table-R1, a three-stage RL framework that enhances multimodal table understanding through: (1) Warm-up that prompts initial perception and reasoning capabilities, (2) Perception Alignment GRPO (PA-GRPO), which employs continuous Tree-Edit-Distance Similarity (TEDS) rewards for recognizing table structures and contents, and (3) Hint-Completion GRPO (HC-GRPO), which utilizes fine-grained rewards of residual steps based on the hint-guided question. Extensive experiments demonstrate that Table-R1 can boost the model's table reasoning performance obviously on both held-in and held-out datasets, outperforming SFT and GRPO largely. Notably, Qwen2-VL-7B with Table-R1 surpasses larger specific table understanding models (e.g., Table-LLaVA 13B), even achieving comparable performance to the closed-source model GPT-4o on held-in datasets, demonstrating the efficacy of each stage of Table-R1 in overcoming initialization bottlenecks and reward sparsity, thereby advancing robust multimodal table understanding.

表格理解强化学习多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。