专精文档元素精确定位,提升复杂文档理解能力
An LMM for Precisely Grounding Elements in Documents

- 构建包含细粒度坐标的高质量合成文档数据集,支持精准定位
- 联合监督强化学习,使定位结果更好支撑推理任务
- 可定位简历中的个人信息等复杂内容,适合文档智能场景
文档中的视觉定位是大型多模态模型在文档理解、深度研究和文档错误检测等领域的关键能力。然而,现有方法在文本密集的文档图像中定位精度较差,难以准确找到可靠推理所需的文档关键元素。为解决这一问题,我们提出PreciseDoc,一种专为精确元素定位设计的LMM,可进一步优化用于文档VQA任务。具体地,通过两条流水线大规模生成带有精细坐标标注的高质量文档,包括带相机效果的合成手写文档,以增强基础定位能力。模型不仅可实现单个文本的定位,还能定位简历中的个人信息等复杂内容。此外,我们引入一种联合监督的强化学习训练范式,同时优化定位与推理,提升定位证据对推理的贡献。在多个基准上的综合评估表明,所提出的数据与方法在文档空间定位与理解方面具有显著优势。
原文摘要 · Abstract (English)
Visual grounding in documents is a crucial ability for Large Multimodal Models (LMMs) in areas such as document understanding, deep research and document error detection. However, existing approaches exhibit poor grounding precision in text-rich document images, often failing to accurately locate the critical document elements needed for reliable reasoning. To address this gap, we introduce PreciseDoc, an LMM specifically designed for precise element grounding and can be further optimized for Document VQA tasks. Specifically, to enhance the basic localization capability, we construct challenging training data by two pipelines capable of mass-producing high-quality documents with paired metadata of fine-grained coordinates, including synthetic hand-filled documents with camera effects. The model develops more real-world functions beyond straightforward localization of single text, such as locating personal information from CVs. Furthermore, we introduce a training paradigm for visual grounded reasoning where the grounding and reasoning are supervised jointly with reinforcement learning to improve the contribution of the grounded evidence. A comprehensive evaluation on various benchmarks demonstrates the advantage of the proposed data and methods in document spatial grounding and document understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。