测试大模型对高棉文文档的理解能力,发现其处理英文和数字效果好,但本地文字仍困难。
Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

- 用高棉文票据图像测试主流多模态模型表现
- 直接推理准确率51.9%,外接OCR可达61.9%
- 对高棉语和混合文字答案理解仍不理想,适合关注低资源语言研究者
近期多模态大模型在文档理解、视觉问答和文本提取方面取得进展,但在低资源非拉丁语环境下的可靠性仍存疑。高棉文文档尤其具有挑战性,因其包含复杂字形、高棉语-英语混写字段及柬币与美元双重货币数值。现有高棉文文档视觉问答资源有限。本文对开源多模态大模型在高棉文文档图像上的表现进行初步诊断评估。基于已有的KH-FUNSD数据集构建评估子集,涵盖发票、收据、报价单等商业表格,问题以英、高棉文提出,答案保留原始英文、高棉文、混合脚本或数值形式。未建立完整公开基准,而是考察现有模型的能力与失败模式。评估代表性的Qwen-VL系列模型,比较直接图像提示、解析器辅助和外部OCR辅助配置,使用Qwen3-VL-8B。直接提示下Qwen3-VL-8B表现最佳,整体准确率达51.9%,但高棉语和混合脚本答案表现仍弱。外部OCR显著提升性能,Tesseract达61.9%,PaddleOCR达61.6%。然而,高棉语答案识别难度远高于英文和数值字段。结果表明当前多模态大模型可处理清晰英文与结构化数值内容,但可靠理解原生高棉文文档仍是开放挑战。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges because they contain complex script forms, mixed Khmer-English fields, and monetary values in both Cambodian Riel and US Dollars. Available resources for Khmer Document VQA are also limited. This paper presents a pilot diagnostic evaluation of open MLLMs on Khmer document images. We construct an evaluation subset from the previously introduced KH-FUNSD collection, covering invoices, receipts, quotations, and other business forms. The subset includes questions in English and Khmer, with answers retained in their original English, Khmer, mixed-script, or numeric forms. Rather than introducing a full public benchmark, this study examines the capabilities and failure modes of existing models. We evaluate representative open Qwen-VL models using direct image-based prompting and compare parser-assisted and external OCR-assisted configurations with Qwen3-VL-8B. Direct Qwen3-VL-8B outperforms smaller models, achieving 51.9% overall accuracy, although performance remains limited for Khmer-script and mixed-script answers. External OCR produces the strongest results, reaching 61.9% with Tesseract and 61.6% with PaddleOCR. Nevertheless, Khmer-script answers remain substantially more difficult than English and numeric fields. The results indicate that current MLLMs can process visually clear English and structured numeric content, but reliable native Khmer document understanding remains an open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。