arXiv:2506.14079cs.AI2025-06Conference of the …被引 1

让智能体在纯图像上自动填表,突破定位与多模态理解瓶颈。

FormGym: Doing Paperwork with Agents

  • 构建纯图像表单填空基准,含55份文档、432个字段和3类任务。
  • 基线模型准确率不足1%,主因是文本位置识别能力差。
  • 提出FieldFinder工具,使所有模型性能最高提升至56%。

填写表单是一项复杂且耗时的任务。在无法使用OCR、排版PDF文本或DOM的纯图像领域中,表单填写尤为困难。对计算机代理而言,这需要多模态理解、信息检索和工具使用等多种能力。我们提出一个全新的表单填写基准,包含55份文档、432个字段及3项任务,每位用户需掌握236个特征。实验发现,基线视觉语言模型(VLAs)在多数情况下准确率低于1%,主要源于定位能力薄弱。图形用户界面(GUI)代理虽表现较好,但准确率仅在10.6%-68.0%之间,且成本高、延迟大。为此,我们还提出了FieldFinder工具,帮助大语言模型精准识别文本填写位置。在六种实验条件下,引入FieldFinder后,所有模型表现均达到或超过原有水平,最高准确率从2%提升至56%。

原文摘要 · Abstract (English)

Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without access to OCR, typeset PDF text, or a DOM. For computer agents, it requires multiple abilities, including multi-modal understanding, information retrieval, and tool-use. We present a novel form-filling benchmark consisting of 432 fields spread across 55 documents and 3 tasks, requiring knowledge of 236 features per user. We find that baseline VLAs achieve less than 1% accuracy in most cases, primarily due to poor localization ability. GUI agents also struggle, scoring between 10.6-68.0% despite high cost and latency. Therefore, we also contribute FieldFinder, a tool to assist LLMs in identifying where to place text on a form. With FieldFinder, all models achieve equal or better performance in all six study conditions, with a maximum increase from 2% to 56%.

智能体表单填写多模态工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。