arXiv:2601.19578cs.CL2026-01被引 1

Yunque框架提升大模型深度研究能力,解决长任务混乱与错误蔓延问题。

Yunque DeepResearch Technical Report

  • 分层模块化设计,用工具池和子代理协同完成复杂任务
  • 动态上下文管理减少信息过载,提升长期任务稳定性
  • 主动监控与纠错机制,适合需要高可靠性的研究应用

深度研究已成为自主智能体的关键能力,使大语言模型能够处理复杂、开放式的任务。然而,其潜力受限于长期任务中上下文噪声加剧、脆弱性导致级联错误以及缺乏模块可扩展性等问题。为此,我们提出Yunque DeepResearch,一个分层、模块化且鲁棒的框架。该架构包含三个核心组件:(1) 中心化的多智能体调度系统,将子任务分配给工具池与专用子代理;(2) 动态上下文管理机制,将已完成的子目标结构化为语义摘要,缓解信息过载;(3) 主动监督模块,通过异常检测与上下文修剪增强系统韧性。Yunque DeepResearch在GAIA、BrowseComp、BrowseComp-ZH及Humanity's Last Exam等多个代理深度研究基准上达到领先性能。我们开源了框架、可复现实现及应用案例,以支持社区发展。

原文摘要 · Abstract (English)

Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full potential is hindered by critical limitations, including escalating contextual noise in long-horizon tasks, fragility leading to cascading errors, and a lack of modular extensibility. To address these challenges, we introduce Yunque DeepResearch, a hierarchical, modular, and robust framework. The architecture is characterized by three key components: (1) a centralized Multi-Agent Orchestration System that routes subtasks to an Atomic Capability Pool of tools and specialized sub-agents; (2) a Dynamic Context Management mechanism that structures completed sub-goals into semantic summaries to mitigate information overload; and (3) a proactive Supervisor Module that ensures resilience through active anomaly detection and context pruning. Yunque DeepResearch achieves state-of-the-art performance across a range of agentic deep research benchmarks, including GAIA, BrowseComp, BrowseComp-ZH, and Humanity's Last Exam. We open-source the framework, reproducible implementations, and application cases to empower the community.

深度研究多智能体上下文管理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。