arXiv:2607.21780cs.CLcs.AI2026-07

首个面向孟加拉语表单的多模态文档拆分基准,解决低资源语言文档重组难题。

Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms

论文配图:Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
图 1 · 摘自论文原文
  • 基于图像的双语(孟加拉语-英语)文档包拆分,直接处理页面图像。
  • 模型在页面聚类上表现尚可,但乱序后恢复原始顺序能力差。
  • 揭示页面排列是核心挑战,适用于低资源文档理解研究者。

文档包(多个文档拼接成单一文件)在政府与行政流程中常见,但将其拆分为独立文档极具挑战性,尤其对低资源语言而言。本文提出 Khondo(孟加拉语意为‘分割’),首个针对孟加拉国政府表单的文档包拆分基准。不同于以往基于英文和OCR文本的数据集,Khondo 为双语(孟加拉语-英语)且视觉原生——模型直接操作页面图像。数据涵盖五种拼接方式(从顺序到完全打乱),覆盖14个行政领域,提供真实边界、领域类型和页序标注。零样本评估显示,多模态大模型虽能较好聚类页面至源文档,但在完全打乱后难以恢复原始页序。通过两项受控分析发现:(a)明确的页序指令虽必要但不足;(b)英文文档包的排序更可靠,表明页序重建是主要瓶颈,语言为次要但持续影响因素。Khondo 确立了视觉驱动的低资源文档理解中页序重建这一关键开放问题,并提供了可控基准以衡量进展。数据集与代码已公开于 https://huggingface.co/datasets/Mausul/khondo。

原文摘要 · Abstract (English)

Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is difficult, especially for low-resource languages. We introduce Khondo (Bangla for split/segment), the first benchmark for document packet splitting on Bangladeshi government forms. Unlike prior English and OCR-text-based datasets, Khondo is bilingual (Bangla--English) and vision-native; where models operate directly on page images. It spans five concatenation schemes, from sequential to fully shuffled, across 14 administrative domains, with ground-truth boundaries, domain types, and page order. Zero-shot evaluation of MLLMs shows they cluster pages into their source documents fairly well but struggle in restoring the original page order once shuffled. To isolate what drives this difficulty, we run two controlled analyses, varying the prompt instruction and then the packet language. Both primarily affect ordering rather than clustering: (a) explicit page-order instructions are necessary but insufficient, and (b) English packets are ordered more reliably than Bangla, making page arrangement the dominant challenge and language a secondary but consistent factor. Khondo establishes page-order reconstruction as a key open problem in vision-based, low-resource document understanding, and provides a controlled benchmark for measuring progress toward solving it. Our dataset and code is available at https://huggingface.co/datasets/Mausul/khondo

文档理解多模态低资源图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。