arXiv:2501.02235cs.CL2025-01综述被引 5

系统梳理图文文档问答的现状与挑战,指明未来方向。

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

  • 归纳当前图文文档理解的主流方法
  • 指出处理流程中缺乏共识的关键难题
  • 适合关注文档智能的科研与工程人员

视觉丰富文档理解领域涉及对扫描或数字生成的图文混合文档进行交互式处理,正快速发展但仍缺乏关键处理流程的共识。本文全面综述了最新技术,强调其优势与局限,揭示领域核心挑战,并提出有前景的研究方向。

原文摘要 · Abstract (English)

The field of visually-rich document understanding, which involves interacting with visually-rich documents (whether scanned or born-digital), is rapidly evolving and still lacks consensus on several key aspects of the processing pipeline. In this work, we provide a comprehensive overview of state-of-the-art approaches, emphasizing their strengths and limitations, pointing out the main challenges in the field, and proposing promising research directions.

文档理解问答系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。