用多智能体协作与动态计算分配,提升文档理解的准确性与推理能力。
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
- 拆分文档处理为规划、执行、判断、回答四类智能体协同工作。
- 在多个基准上以更小模型实现9.9%-11.5%性能提升。
- 适合需要高准确性和复杂推理的文档分析场景。
视觉语言模型的单体规模扩展在文档理解与推理任务中已陷入瓶颈,难以应对文档特有的流程化推理、认知复杂性与事实准确性需求。为此,我们提出MACT——一种基于智能体自适应测试时缩放的多智能体协作框架,开创了流程化扩展的新范式。MACT将视觉文档处理流程分解为规划、执行、判断和答案四个专业化智能体,缓解认知过载,并引入关键的事实校正机制。该协作架构通过智能体级自适应测试时缩放策略,根据各功能模块的复杂度与冗余程度动态分配计算资源。在多个视觉文档理解基准上评估,MACT以更小参数量取得优异表现,能有效适应各类文档场景,且不损害其通用或数学推理能力。三种MACT变体平均性能均位列前三,相比基线模型提升9.9%-11.5%。源代码将公开发布。
原文摘要 · Abstract (English)
The dominant paradigm of monolithic scaling in Vision-Language Models (VLMs) is failing for understanding and reasoning in documents, yielding diminishing returns as it struggles with the inherent need of this domain for document-based procedural reasoning, cognitive complexity, and factual accuracy. To this end, we introduce MACT, a Multi-Agent Collaboration framework with agent-wise adaptive Test-time scaling that pioneers a paradigm shift to procedural scaling, adapting dynamically to the functional entities of visual documents understanding and reasoning. MACT decomposes the visual document processing flow into four specialized agents, i.e., planning, execution, judgment, and answer, to resolve cognitive overload and introduce a critical self-correction loop for factual grounding. This collaborative architecture is amplified by an agent-wise adaptive test-time scaling strategy that intelligently allocates computational resources based on the complexity and redundancy of each functionality. Evaluated on multiple visual document understanding benchmarks, MACT achieves superior performance with a smaller parameter scale, adapting effectively to various document scenarios without compromising its general or mathematical reasoning capabilities. The three variants of MACT consistently attain top-three average performance rankings, with average performance enhancements of 9.9-11.5% over the base models. The source code will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。