arXiv:2509.17589cs.AI2025-09NeurIPS被引 7

用强化学习生成高保真LaTeX表格代码,精准还原复杂结构。

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

  • 基于多模态大模型与强化学习,结合结构与视觉双奖励优化生成质量。
  • 在复杂表格上超越现有方法,结构相似度达91.7%,视觉保真度提升显著。
  • 适合需要自动化排版科研论文的作者或期刊编辑使用。

本文研究从表格图像生成高质量LaTeX代码的任务,旨在自动化重建可直接发表的表格。核心挑战在于处理大型、深度嵌套或语义复杂的表格,现有方法常失效。我们通过全面分析识别关键问题,并指出当前评估协议的局限性。为此提出基于强化学习的多模态大语言模型框架:在大规模表格-TeX数据集上微调预训练多模态模型,并引入基于组相对策略优化(GRPO)的双奖励机制,同时优化输出的结构正确性与渲染后的视觉保真度。采用结合TEDS-Structure与CW-SSIM的混合评估方案,实验表明该方法在结构复杂表格上表现最优,结构相似度达91.7%,验证了其有效性与鲁棒性。

原文摘要 · Abstract (English)

In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables -- those with large sizes, deeply nested structures, and semantically rich or irregular cell content -- where existing methods often fail. We begin with a comprehensive analysis, identifying key challenges and highlighting the limitations of current evaluation protocols. To overcome these issues, we propose a reinforced multimodal large language model (MLLM) framework, where a pre-trained MLLM is fine-tuned on a large-scale table-to-LaTeX dataset. To further improve generation quality, we introduce a dual-reward reinforcement learning strategy based on Group Relative Policy Optimization (GRPO). Unlike standard approaches that optimize purely over text outputs, our method incorporates both a structure-level reward on LaTeX code and a visual fidelity reward computed from rendered outputs, enabling direct optimization of the visual output quality. We adopt a hybrid evaluation protocol combining TEDS-Structure and CW-SSIM, and show that our method achieves state-of-the-art performance, particularly on structurally complex tables, demonstrating the effectiveness and robustness of our approach.

表格生成强化学习LaTeX多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。