arXiv:2603.00707cs.CV2026-03被引 1

首个针对高棉语场景文档布局检测的研究,解决数据少、文字复杂难题。

Towards Khmer Scene Document Layout Detection

  • 构建首个高棉语场景文档布局专用数据集与增强工具
  • 提出基于YOLO和旋转框的布局检测模型,应对透视变形
  • 适合高棉语文档分析研究者及多语言场景识别开发者

尽管拉丁文文档版面分析因大型多模态模型的发展取得显著进展,但高棉语由于标注训练数据稀缺,相关研究仍受限,尤其在存在视角扭曲和复杂背景的场景文档中更为突出。高棉文字结构复杂,包含变音符号和多层字符叠加,现有基于拉丁文的版面分析模型难以准确划分语义版面单元,尤其是在密集文本区域(如列表项)表现不佳。本文首次系统研究高棉语场景文档布局检测,贡献三个核心部分:(1) 面向高棉语场景布局的鲁棒训练与评测数据集;(2) 开源文档增强工具,可生成真实感强的合成场景文档以扩充训练数据;(3) 基于YOLO架构并采用旋转边界框(OBB)的布局检测基线模型,有效处理几何畸变。为推动高棉语文档分析与识别(DAR)领域发展,我们将在审查中的受控仓库中公开模型、代码与数据集。

原文摘要 · Abstract (English)

While document layout analysis for Latin scripts has advanced significantly, driven by the advent of large multimodal models (LMMs), progress for the Khmer language remains constrained because of the scarcity of annotated training data. This gap is particularly acute for scene documents, where perspective distortions and complex backgrounds challenge traditional methods. Given the structural complexities of Khmer script, such as diacritics and multi-layer character stacking, existing Latin-based layout analysis models fail to accurately delineate semantic layout units, particularly for dense text regions (e.g., list items). In this paper, we present the first comprehensive study on Khmer scene document layout detection. We contribute a novel framework comprising three key elements: (1) a robust training and benchmarking dataset specifically for Khmer scene layouts; (2) an open-source document augmentation tool capable of synthesizing realistic scene documents to scale training data; and (3) layout detection baselines utilizing YOLO-based architectures with oriented bounding boxes (OBB) to handle geometric distortions. To foster further research in the Khmer document analysis and recognition (DAR) community, we release our models, code, and datasets in this gated repository (in review).

文档分析高棉语场景文档版面检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。