arXiv:2601.18267cs.IR2026-01

用智能代理协作提升企业级问答的深度与可信度

Orchestrating Specialized Agents for Trustworthy Enterprise RAG

  • 采用中心协调的多代理迭代式调查,替代传统单轮检索生成
  • 在DeepResearch基准上达52.65分,优于商用系统77.2%偏好胜率
  • 支持可追溯证据链与长文本压缩,适合高可靠性场景使用

检索增强生成(RAG)在企业知识工作中潜力巨大,但在需要深度综合、严格可追溯性和对模糊提示容错的高风险决策场景中表现不佳。单轮检索-生成流程常导致浅层摘要、证据不一致和完整性验证薄弱。本文提出ADORE(面向企业研究的自适应深度协同框架),以中央协调器和多个专业化代理组成的架构取代线性流程。其核心在于结构化记忆库(带显式论点-证据关联与章节级可接受证据),实现可追溯报告生成与系统性完整性检查。贡献包括:(1) 记忆锁定合成——报告生成仅限于带有章节级可接受证据的结构化记忆库(论点-证据图),确保可追溯与有据可依;(2) 证据覆盖引导执行——通过检索-反思循环审计章节级证据覆盖率,触发针对性补充检索,并基于证据驱动停止准则终止;(3) 章节级长上下文接地——通过章节打包、剪枝与引用保留压缩,克服上下文长度限制,实现长文本合成。在评估套件中,ADORE在DeepResearch Bench上排名第一(52.65分),在DeepConsult上以77.2%的头对头偏好胜率超越商用系统。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict traceability, and recovery from underspecified prompts. One-pass retrieval-and-write pipelines frequently yield shallow summaries, inconsistent grounding, and weak mechanisms for completeness verification. We introduce ADORE (Adaptive Deep Orchestration for Research in Enterprise), an agentic framework that replaces linear retrieval with iterative, user-steered investigation coordinated by a central orchestrator and a set of specialized agents. ADORE's key insight is that a structured Memory Bank (a curated evidence store with explicit claim-evidence linkage and section-level admissible evidence) enables traceable report generation and systematic checks for evidence completeness. Our contributions are threefold: (1) Memory-locked synthesis - report generation is constrained to a structured Memory Bank (Claim-Evidence Graph) with section-level admissible evidence, enabling traceable claims and grounded citations; (2) Evidence-coverage-guided execution - a retrieval-reflection loop audits section-level evidence coverage to trigger targeted follow-up retrieval and terminates via an evidence-driven stopping criterion; (3) Section-packed long-context grounding - section-level packing, pruning, and citation-preserving compression make long-form synthesis feasible under context limits. Across our evaluation suite, ADORE ranks first on DeepResearch Bench (52.65) and achieves the highest head-to-head preference win rate on DeepConsult (77.2%) against commercial systems.

企业问答智能代理可信生成RAG优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。