arXiv:2512.11849cs.CLcs.AI2025-12被引 2

首个高精度标注的高棉语商业文档数据集,助力低资源语言文档理解

KH-FUNSD: A Hierarchical and Fine-Grained Layout Analysis Dataset for Low-Resource Khmer Business Document

  • 构建三级标注体系:区域划分、实体识别与细粒度语义分类
  • 涵盖收据、发票等170+份文档,支持字段级信息提取
  • 为非拉丁语系低资源语言研究提供首个基准数据集

低资源非拉丁文字的自动化文档版面分析仍面临重大挑战。高棉语是柬埔寨超过1700万人日常使用的语言,在文档人工智能工具开发中却鲜受关注。商业文档资源尤为匮乏,而这正是公共管理与私营企业运作的关键。为此,我们提出首个公开可用的高棉语表单文档理解数据集——KH-FUNSD,包含收据、发票和报价单等类型。其标注框架采用三级设计:(1)区域检测,将文档划分为标题、表单域、页脚等核心区域;(2)类似FUNSD的标注,区分问题、答案、标题及其他关键实体及其关系;(3)细粒度分类,赋予具体语义角色如字段标签、数值、标题、页脚、符号等。该多层级方法支持全面版面分析与精准信息抽取。我们对多个主流模型进行基准测试,首次提供高棉语商业文档的基线结果,并讨论非拉丁语系、低资源语言带来的独特挑战。KH-FUNSD数据集及文档将公开发布。

原文摘要 · Abstract (English)

Automated document layout analysis remains a major challenge for low-resource, non-Latin scripts. Khmer is a language spoken daily by over 17 million people in Cambodia, receiving little attention in the development of document AI tools. The lack of dedicated resources is particularly acute for business documents, which are critical for both public administration and private enterprise. To address this gap, we present \textbf{KH-FUNSD}, the first publicly available, hierarchically annotated dataset for Khmer form document understanding, including receipts, invoices, and quotations. Our annotation framework features a three-level design: (1) region detection that divides each document into core zones such as header, form field, and footer; (2) FUNSD-style annotation that distinguishes questions, answers, headers, and other key entities, together with their relationships; and (3) fine-grained classification that assigns specific semantic roles, such as field labels, values, headers, footers, and symbols. This multi-level approach supports both comprehensive layout analysis and precise information extraction. We benchmark several leading models, providing the first set of baseline results for Khmer business documents, and discuss the distinct challenges posed by non-Latin, low-resource scripts. The KH-FUNSD dataset and documentation will be available at URL.

文档理解低资源语言高棉语版面分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。