用协作智能体分解复杂文档问答任务,提升推理准确性。
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
- 设计多智能体协同框架,按任务类型动态调用专用处理模块。
- 在三个基准上超越现有方法,显著提升复杂问题解答能力。
- 适合需要精细推理与跨模态理解的文档分析场景。
文档视觉问答(DocVQA)对现有视觉语言模型仍具挑战性,尤其在复杂推理与多步流程中。当前方法难以将复杂问题分解为可管理子任务,且无法有效利用不同文档元素的专用处理路径。本文提出ORCA:面向文档视觉问答的协作式推理框架,通过智能体协同与迭代优化解决上述问题。系统首先由推理智能体将问题分解为逻辑步骤,再通过路由机制激活来自专用智能体库的任务特定智能体。框架配备多个专注不同模态的智能体,实现对多样化文档成分的细粒度理解与协同推理。为确保答案可靠性,引入辩论机制与压力测试,并在必要时采用正反方裁决流程,最后由健康检查器保障输出格式一致性。在三个基准上的实验表明,该方法显著优于现有最先进水平,确立了视觉语言推理中协作智能体系统的全新范式。
原文摘要 · Abstract (English)
Document Visual Question Answering (DocVQA) remains challenging for existing Vision-Language Models (VLMs), especially under complex reasoning and multi-step workflows. Current approaches struggle to decompose intricate questions into manageable sub-tasks and often fail to leverage specialized processing paths for different document elements. We present ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering, a novel multi-agent framework that addresses these limitations through strategic agent coordination and iterative refinement. ORCA begins with a reasoning agent that decomposes queries into logical steps, followed by a routing mechanism that activates task-specific agents from a specialized agent dock. Our framework leverages a set of specialized AI agents, each dedicated to a distinct modality, enabling fine-grained understanding and collaborative reasoning across diverse document components. To ensure answer reliability, ORCA employs a debate mechanism with stress-testing, and when necessary, a thesis-antithesis adjudication process. This is followed by a sanity checker to ensure format consistency. Extensive experiments on three benchmarks demonstrate that our approach achieves significant improvements over state-of-the-art methods, establishing a new paradigm for collaborative agent systems in vision-language reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。