用强化学习生成高保真LaTeX表格代码,精准还原复杂结构。
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
- 基于多模态大模型与强化学习,结合结构与视觉双奖励优化生成质量。
- 在复杂表格上超越现有方法,结构相似度达91.7%,视觉保真度提升显著。
- 适合需要自动化排版科研论文的作者或期刊编辑使用。
本文研究从表格图像生成高质量LaTeX代码的任务,旨在自动化重建可直接发表的表格。核心挑战在于处理大型、深度嵌套或语义复杂的表格,现有方法常失效。我们通过全面分析识别关键问题,并指出当前评估协议的局限性。为此提出基于强化学习的多模态大语言模型框架:在大规模表格-TeX数据集上微调预训练多模态模型,并引入基于组相对策略优化(GRPO)的双奖励机制,同时优化输出的结构正确性与渲染后的视觉保真度。采用结合TEDS-Structure与CW-SSIM的混合评估方案,实验表明该方法在结构复杂表格上表现最优,结构相似度达91.7%,验证了其有效性与鲁棒性。
原文摘要 · Abstract (English)
In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables -- those with large sizes, deeply nested structures, and semantically rich or irregular cell content -- where existing methods often fail. We begin with a comprehensive analysis, identifying key challenges and highlighting the limitations of current evaluation protocols. To overcome these issues, we propose a reinforced multimodal large language model (MLLM) framework, where a pre-trained MLLM is fine-tuned on a large-scale table-to-LaTeX dataset. To further improve generation quality, we introduce a dual-reward reinforcement learning strategy based on Group Relative Policy Optimization (GRPO). Unlike standard approaches that optimize purely over text outputs, our method incorporates both a structure-level reward on LaTeX code and a visual fidelity reward computed from rendered outputs, enabling direct optimization of the visual output quality. We adopt a hybrid evaluation protocol combining TEDS-Structure and CW-SSIM, and show that our method achieves state-of-the-art performance, particularly on structurally complex tables, demonstrating the effectiveness and robustness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。