arXiv:2504.10419cs.CL2025-04被引 1

构建首个专门评测大模型识别复选框的基准数据集,解决文档理解中的关键盲区。

Unchecked and Overlooked: Addressing the Checkbox Blind Spot in Large Language Models with CheckboxQA

  • 设计针对复选框识别的任务评估集CheckboxQA
  • 实测主流大模型在复选框任务上准确率普遍低于60%
  • 适合法律、金融等对文档细节敏感领域的研究与应用

复选框在真实文档处理中至关重要,其勾选状态直接影响信息提取与决策。尽管大视觉语言模型在众多任务上表现优异,但在解析可勾选内容方面仍存在显著缺陷。这一问题在法律、金融等行业尤为严重,漏检一个复选框可能导致重大合规或合同风险。为此,我们提出CheckboxQA数据集,专门用于评估和提升模型在复选框相关任务上的能力。该数据集揭示了当前模型的局限性,并为改进文档理解系统提供了重要工具,对法律科技、金融等领域具有深远意义。数据集已公开:https://github.com/Snowflake-Labs/CheckboxQA

原文摘要 · Abstract (English)

Checkboxes are critical in real-world document processing where the presence or absence of ticks directly informs data extraction and decision-making processes. Yet, despite the strong performance of Large Vision and Language Models across a wide range of tasks, they struggle with interpreting checkable content. This challenge becomes particularly pressing in industries where a single overlooked checkbox may lead to costly regulatory or contractual oversights. To address this gap, we introduce the CheckboxQA dataset, a targeted resource designed to evaluate and improve model performance on checkbox-related tasks. It reveals the limitations of current models and serves as a valuable tool for advancing document comprehension systems, with significant implications for applications in sectors such as legal tech and finance. The dataset is publicly available at: https://github.com/Snowflake-Labs/CheckboxQA

文档理解大模型评估法律科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。