将多种信息搜索代理合并为统一模型,显著降低训练成本并保持跨领域能力。
Exploring Information Seeking Agent Consolidation
- 通过参数级融合方式整合多个专用代理模型
- 仅需数据混合1/10训练成本即可达到相当性能
- 适合追求高效、稳定跨领域部署的研究者
信息搜索代理已成为知识密集型任务的重要范式,但现有系统仍局限于开放网络、文档或本地知识库,难以实现规模化和跨领域部署。本文首次系统性地研究将这些信息搜索代理整合为单一基础代理模型的可行性。比较了两种范式:数据级混合(在混合数据集上训练统一模型)与参数级合并(在参数空间中融合独立训练的专家模型),覆盖3种训练场景,并在10个基准测试上评估了26种代表性参数级方法。为在异构基准间公平比较,提出几何综合得分与不平衡得分以衡量整体表现与任务偏差。分析表明:(i) 设计合理的参数级合并可在远低于数据混合的训练成本下实现同等性能,且对任务顺序不敏感;(ii) 参数级合并能结构性保留数据混合会遗忘的跨域能力;(iii) 跨场景稳定性强烈依赖于合并质量。基于观察结果,提炼出方法选择指南与下一代合并算子的设计原则。
原文摘要 · Abstract (English)
Information-seeking agents have emerged as a powerful paradigm for knowledge-intensive tasks, yet today's systems remain specialized for the open web, documents, or local knowledge bases, hindering scalable and cross-domain deployment. We present the first systematic empirical study of consolidating these information-seeking agents into a single foundation agentic model. We compare two paradigms -- \emph{data-level mixing}, which trains a unified model on a mixture of datasets, and \emph{parameter-level merging}, which merges independently trained experts in parameter space -- across 3 training scenarios, evaluating \textbf{26} representative parameter-level methods on \textbf{10} benchmarks. To compare across heterogeneous benchmarks, we introduce a geometric Composite Score and an Imbalance Score that describe overall performance and task skew. Our analysis shows that (i) well-designed parameter-level merging attains parity with data mixing at a fraction of its training cost and is order-agnostic; (ii) parameter-level merging structurally preserves out-of-domain capabilities that data mixing universally forgets; and (iii) cross-scenario stability is strongly tied to consolidation quality. We distil our observations into a method-selection guide and design principles for next-generation merging operators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。