arXiv:2510.13366cs.CLcs.AI2025-10综述被引 7

大模型重塑文档智能,带来理解与生成的飞跃。

Document Intelligence in the Era of Large Language Models: A Survey

  • 用单向解码器大模型替代传统架构,提升文档理解能力。
  • 在多模态、多语言和检索增强场景中实现显著性能提升。
  • 适合研究者与从业者了解文档智能前沿进展。

文档人工智能(DAI)已成为关键应用领域,其发展深受大型语言模型(LLMs)推动。早期方法依赖编码器-解码器架构,而仅使用解码器的LLMs已彻底改变DAI,显著提升文档的理解与生成能力。本文综述了DAI的演进历程,重点分析当前在多模态、多语言及检索增强型文档智能方面的研究进展与挑战,并提出未来方向,包括基于代理的方法和面向文档的通用模型。本研究旨在系统梳理文档智能的最新成果及其对学术与实际应用的影响。

原文摘要 · Abstract (English)

Document AI (DAI) has emerged as a vital application area, and is significantly transformed by the advent of large language models (LLMs). While earlier approaches relied on encoder-decoder architectures, decoder-only LLMs have revolutionized DAI, bringing remarkable advancements in understanding and generation. This survey provides a comprehensive overview of DAI's evolution, highlighting current research attempts and future prospects of LLMs in this field. We explore key advancements and challenges in multimodal, multilingual, and retrieval-augmented DAI, while also suggesting future research directions, including agent-based approaches and document-specific foundation models. This paper aims to provide a structured analysis of the state-of-the-art in DAI and its implications for both academic and practical applications.

文档智能大模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。