arXiv:2503.11576cs.CV2025-03ICCV被引 55

256M参数模型实现文档元素端到端精准转换。

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

  • 用新标记格式DocTags统一处理整页文档内容与位置。
  • 在256M参数下性能媲美大27倍的模型,支持多类型文档。
  • 适合需要轻量级、高精度文档转换的开发者和研究者。

我们提出SmolDocling,一个超紧凑的视觉语言模型,用于端到端文档转换。该模型通过生成一种名为DocTags的新通用标记格式,全面处理整页内容,精确捕捉所有页面元素及其空间位置。与依赖大型基础模型或复杂手工管道的现有方法不同,SmolDocling在仅256M参数下实现了对代码块、表格、公式、图表、列表等文档元素的内容、结构和空间位置的准确还原,覆盖商业文件、学术论文、技术报告、专利和表单等多种类型,显著超越了以往以科学论文为主的局限。此外,我们还贡献了公开可获取的图表、表格、公式和代码识别数据集。实验表明,其性能可与参数量高达27倍大的视觉语言模型相媲美,同时大幅降低计算开销。模型已可使用,数据集即将公开。

原文摘要 · Abstract (English)

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon.

文档转换轻量模型多模态视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。