SlideAgent用分层智能体解析多页幻灯片,让模型更懂图文布局和跨页关联。
SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
- 分三层(全局、页面、元素)推理,构建结构化理解框架
- 在多页文档任务中,相比开源模型提升9.8%准确率
- 适合需要细粒度分析幻灯片的科研与办公场景
多页视觉文档如手册、宣传册、演示文稿和海报通过版式、色彩、图标和跨页引用传递关键信息。尽管多模态大语言模型(MLLMs)为文档理解提供了新可能,现有系统在处理复杂多页视觉文档时仍面临挑战,尤其在元素与页面间的细粒度推理方面表现不足。我们提出SlideAgent,一种适用于多模态、多页、多版式文档(特别是幻灯片)的通用智能体框架。SlideAgent采用专用智能体,将推理分解为三个层次——全局、页面和元素,构建查询无关的结构化表示,同时捕捉整体主题与细节视觉或文本线索。推理过程中,SlideAgent动态激活对应层级的智能体,整合输出生成连贯且上下文感知的答案。大量实验表明,SlideAgent在准确率上显著优于专有模型(+7.9%)和开源模型(+9.8%)。
原文摘要 · Abstract (English)
Multi-page visual documents such as manuals, brochures, presentations, and posters convey key information through layout, colors, icons, and cross-slide references. While multimodal large language models (MLLMs) offer opportunities in document understanding, current systems struggle with complex, multi-page visual documents, particularly in fine-grained reasoning over elements and pages. We introduce SlideAgent, a versatile agentic framework for understanding multi-modal, multi-page, and multi-layout documents, especially slide decks. SlideAgent employs specialized agents and decomposes reasoning into three specialized levels--global, page, and element--to construct a structured, query-agnostic representation that captures both overarching themes and detailed visual or textual cues. During inference, SlideAgent selectively activates specialized agents for multi-level reasoning and integrates their outputs into coherent, context-aware answers. Extensive experiments show that SlideAgent significantly improves accuracy over both proprietary (+7.9%) and open-source models (+9.8%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。