构建交互式评测套件,推动多模态表单自动填充研究
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
- 设计交互式评测平台,包含网页界面与真实场景数据集
- 现有最强模型准确率不足5%,暴露布局理解与字段对齐短板
- 适合关注GUI自动化、多模态推理的开发者与研究者
在线表单填写是一项常见但耗时的任务,需大量键盘和鼠标操作。尽管长期以来期望实现‘一键完成’,但现有工具仍以规则为基础,缺乏可泛化的生成能力。多模态大语言模型(MLLM)虽在通用任务中表现良好,但在表单填写中面临布局灵活、文本指令与屏幕字段难以对齐等挑战。为此,我们正式定义表单填写任务,提出FormFactory——一个包含网页界面、后端评估模块和精心构建数据集的交互式评测套件。该基准覆盖多样真实场景,包含多种字段格式,并模拟高保真表单交互。我们对主流MLLM进行了全面评估,发现无一模型准确率超过5%,凸显任务难度。结果也揭示当前模型在视觉布局推理和字段-值对齐方面存在显著局限。我们希望本基准能成为推动稳健实用表单填充智能体研究的基石。
原文摘要 · Abstract (English)
Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one click", existing tools remain largely rule-based and lack generalizable, generative capabilities. Recent advances in Multimodal Large Language Models (MLLMs) have enabled promising agents for GUI-related tasks in general-purpose scenarios. However, they struggle with the unique challenges of form filling, such as flexible layouts and the difficulty of aligning textual instructions with on-screen fields. To bridge this gap, we formally define the form-filling task and propose FormFactory, an interactive benchmarking suite comprising a web-based interface, backend evaluation module, and carefully constructed dataset. Our benchmark covers diverse real-world scenarios, incorporates various field formats, and simulates high-fidelity form interactions. We conduct a comprehensive evaluation of state-of-the-art MLLMs and observe that no model surpasses 5% accuracy, underscoring the inherent difficulty of the task. These findings also reveal significant limitations in current models' visual layout reasoning and field-value alignment abilities. We hope our benchmark can serve as a stepping stone for further research into robust, practical form-filling agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。