arXiv:2607.05614cs.CLcs.AI2026-07中稿 · ECCV

构建首个复杂孟加拉文表单理解基准,解决低资源语言文档解析难题。

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

论文配图:BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
图 1 · 摘自论文原文
  • 设计26类细粒度实体标注体系,涵盖政府多领域复杂表单结构
  • 在零样本与链式思考提示下,主流大模型对细粒度实体定位准确率不足50%
  • 适用于低资源语言文档理解研究,尤其适合多模态大模型评估

文档理解是多模态大语言模型中极具挑战性且影响深远的任务,尤其在面向人类的现实应用中日益重要。然而,由于高质量标注数据稀缺,这类系统在低资源语言如孟加拉语中的应用受限。为此,我们提出BaFCo,一个专注于文档布局分析(DLA)与关键信息提取(KIE)的孟加拉文表单理解基准。该数据集包含200个来自农业、教育、银行和土地管理等领域的多页复杂孟加拉国政府表单。为精准捕捉表单的结构与上下文复杂性,我们定义了包含26种实体类型的细粒度标注方案,以及5种粗粒度实体类别。我们在ChatGPT、Gemini、Claude、Qwen和Kimi系列最新模型上,使用零样本和链式思考提示,在低与高推理设置下进行评估。结果表明,当前多模态大模型在理解孟加拉文表单方面存在明显局限,尤其是在定位高度细粒度实体时表现不佳。数据集与代码已公开于:https://huggingface.co/datasets/Mausul/bafco。

原文摘要 · Abstract (English)

Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limited for low-resource languages such as Bangla due to the scarcity of high-quality annotated data. To address this gap, we introduce BaFCo, a benchmark dataset for Bangla form comprehension with a focus on Document Layout Analysis (DLA) and Key Information Extraction (KIE). BaFCo curates 200 multi-page complex Bangladeshi government forms, sourced from across diverse sectors including agriculture, education, banking, and land management. To accurately capture the structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, along with a separate coarse form entity set consisting of 5 types. We evaluate the latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimi series using zero-shot and chain-of-thought prompts under both low and high reasoning setups. Our results reveal limitations in current MLLMs' ability in comprehending Bangla forms, particularly in accurately localizing highly granular form entities. Our dataset and code is available at: https://huggingface.co/datasets/Mausul/bafco

文档理解低资源语言多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。