通过注意力引导的迭代修正,提升数学公式图像转LaTeX的准确率。
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
- 引入注意力定位与迭代优化,让模型自我纠错
- 在新数据集上性能显著提升,多轮推理效果更优
- 适合需要高精度公式转换的研究与教育场景
Img2LaTeX 是一项重要的实用任务,旨在将图像中的数学表达式和结构化视觉内容转换为LaTeX代码。近年来,视觉语言模型(VLMs)在诸多视觉理解任务中取得显著进展,主要得益于其强大的泛化能力。然而,尽管已有尝试将VLMs应用于Img2LaTeX任务,其表现仍不理想。实证表明,VLMs在细粒度视觉元素(如数学表达式的下标、上标)上易出错,导致LaTeX生成不准确。为此,本文提出 $A^2R^2$:通过注意力引导的视觉推理与迭代修正框架,增强VLM在该任务中的自校正能力,逐步提升生成质量。为实现有效评估,我们构建了新数据集 Img2LaTeX-Hard-1K,包含1,100个精心筛选的挑战性样本。实验结果表明:(1) $A^2R^2$ 在文本与视觉层面多项指标上均有显著提升;(2) 增加推理轮次带来明显性能增益,凸显其在测试时扩展中的潜力;(3) 消融实验与进一步评估验证了方法有效性及核心组件间的协同作用。
原文摘要 · Abstract (English)
Img2LaTeX is a practically important task that involves translating mathematical expressions and structured visual content from images into LaTeX code. In recent years, vision-language models (VLMs) have achieved remarkable progress across a range of visual understanding tasks, largely due to their strong generalization capabilities. However, despite initial efforts to apply VLMs to the Img2LaTeX task, their performance remains suboptimal. Empirical evidence shows that VLMs can be challenged by fine-grained visual elements, such as subscripts and superscripts in mathematical expressions, which results in inaccurate LaTeX generation. To address this challenge, we propose $A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement, a framework that effectively integrates attention localization and iterative refinement within a visual reasoning framework, enabling VLMs to perform self-correction and progressively improve LaTeX generation quality. For effective evaluation, we introduce a new dataset, Img2LaTex-Hard-1K, consisting of 1,100 carefully curated and challenging examples designed to rigorously evaluate the capabilities of VLMs within this task domain. Extensive experimental results demonstrate that: (1) $A^2R^2$ significantly improves model performance across various evaluation metrics spanning both textual and visual levels; (2) Increasing the number of inference rounds yields notable performance gains, underscoring the potential of $A^2R^2$ in test-time scaling scenarios; (3) Ablation studies and further evaluations confirm the effectiveness of our approach and the synergy of its core components during inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。