arXiv:2603.15409cs.CL2026-03中稿 · CVPR

构建东南亚多语言文档与场景文字理解基准,覆盖11种语言

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

  • 设计混合标注流程,结合自动化筛选与本地人验证
  • 包含1.5万页文档和7496个问答对,覆盖多种任务类型
  • 揭示低资源语言模型性能显著下降,推动跨语言研究

多语言文档与场景文字理解在搜索、金融和公共服务中至关重要,但现有基准多聚焦高资源语言,难以评估真实多语言环境下的表现。东南亚语言多样、书写系统复杂、文档类型繁多,挑战更大。我们提出SEA-Vision,一个联合评估文档解析与文本中心视觉问答(TEC-VQA)的基准,涵盖11种东南亚语言。该数据集包含15,234页文档,来自9类代表性文档类型,标注了层级化的页面、块和行级标签;同时提供7,496个TEC-VQA问答对,涵盖文本识别、数值计算、比较分析、逻辑推理和空间理解等能力。为实现高效高质量的多任务标注,我们设计了混合标注流程:结合自动化过滤与评分、基于多模态大模型(MLLM)辅助标注及轻量级母语者验证,大幅降低人工成本并保证质量。我们在多个领先多模态模型上进行评估,发现其在低资源东南亚语言上的性能显著下降,凸显当前技术在跨语言文档理解中的巨大差距。SEA-Vision有望推动全球文档与场景文字理解的发展。

原文摘要 · Abstract (English)

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in realistic multilingual environments. In Southeast Asia, the diversity of languages, complex writing systems, and highly varied document types make this challenge even greater. We introduce SEA-Vision, a benchmark that jointly evaluates Document Parsing and Text-Centric Visual Question Answering (TEC-VQA) across 11 Southeast Asian languages. SEA-Vision contains 15,234 document parsing pages from nine representative document types, annotated with hierarchical page-, block-, and line-level labels. It also provides 7,496 TEC-VQA question-answer pairs that probe text recognition, numerical calculation, comparative analysis, logical reasoning, and spatial understanding. To make such multilingual, multi-task annotation feasible, we design a hybrid pipeline for Document Parsing and TEC-VQA. It combines automated filtering and scoring with MLLM-assisted labeling and lightweight native-speaker verification, greatly reducing manual labeling while maintaining high quality. We evaluate several leading multimodal models and observe pronounced performance degradation on low-resource Southeast Asian languages, highlighting substantial remaining gaps in multilingual document and scene text understanding. We believe SEA-Vision will help drive global progress in document and scene text understanding.

多语言文档理解视觉问答东南亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。