arXiv:2605.10550cs.CL2026-05

首个支持多领域多模态的文档分类基准,模拟真实企业文档复杂结构。

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

  • 构建五层深度层级标签体系,还原企业文档真实组织逻辑。
  • 收录5990份来自12个商业领域的多模态文档,专家标注完整路径。
  • 公开数据集与评测工具,助力工业级文档智能研究。

文档分类是现代企业内容管理的核心,但现有基准仍停留在单一领域、扁平标签的简化范式,难以反映真实业务文档的层次化、多模态和跨领域特性。为弥合这一差距,我们构建了首个多层次、多领域、多模态文档分类基准(MMM-Bench)。该基准包含:(1) 覆盖五个层级的深层层级分类体系,真实体现企业文档组织逻辑;(2) 从阿里巴巴12个商业领域精心采集的5,990份真实多模态文档,均由领域专家手动标注完整层级路径。我们在该基准上建立了涵盖开源模型与API模型的全面基线,并通过系统实验识别出四大核心挑战,提出相应洞见。为推动多层次、跨领域文档分类研究,我们已将全部数据及评估工具开源至 https://github.com/MMMDC-Bench/MMMDC-Bench。

原文摘要 · Abstract (English)

Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear little resemblance to the hierarchical, multi-modal, and cross-domain nature of real-world business documents. This gap not only misrepresents practical complexity but also stifles progress toward industrially viable document intelligence. To bridge this gap, we construct the first Multi-level, Multi-domain, Multi-modal document classification Benchmark (MMM-Bench). MMM-Bench includes (1) a deeply hierarchical taxonomy spanning five levels that capture the authentic organizational logic of business documentation; and (2) 5,990 real-world multi-modal documents meticulously curated from 12 commercial domains in Alibaba. Each document is manually annotated with a complete hierarchical path by domain experts. We establish comprehensive baselines on MMM-Bench, which consists of open-weight models and API-based models. Through systematic experiments, we identify four fundamental challenges within MMM-Bench and propose corresponding insights. To provide a solid foundation for advancing research in multi-level, multi-domain document classification, we release all of the data and the evaluation toolkit of MMM-Bench at https://github.com/MMMDC-Bench/MMMDC-Bench.

文档分类多模态工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。