arXiv:2608.02101cs.CL2026-08

让搜索智能体既专业又通用,打破性能与泛化能力的矛盾

Cross-Domain Hybrid OPD for Generalizable Search Agents

论文配图:Cross-Domain Hybrid OPD for Generalizable Search Agents
图 1 · 摘自论文原文
  • 用跨领域专家知识蒸馏,让专用模型保留通用能力
  • 在真实搜索场景中兼顾高效检索与多任务表现
  • 适合想构建全能型智能搜索系统的开发者

近年来强化学习的进步显著提升了自主搜索代理的能力,使其能够对动态信息源进行复杂规划和迭代检索。然而,为特定搜索行为优化语言模型常带来对齐代价,即搜索性能提升会牺牲通用能力,限制其作为通用助手的效果。本文介绍袁宝搜索代理的训练框架,基于 Hunyuan3 架构,结合代理式强化学习与跨领域专家在线策略蒸馏(OPD)流程。将多个互补通用领域专家的知识蒸馏至搜索专用学生模型中,恢复并进一步增强其广泛能力。不同于将专业化与通用能力视为对立目标,该混合训练策略协同优化二者,有效缓解对齐代价。大量实验表明,该模型在保持竞争性搜索性能的同时,持续提升通用能力,在真实搜索场景中实现了专业化执行与广泛泛化的良好平衡。

原文摘要 · Abstract (English)

Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and iterative retrieval over dynamic information sources. However, optimizing language models for specialized search behaviors often incurs an alignment tax, where gains in search performance come at the expense of general-purpose capabilities, limiting their effectiveness as universal assistants. In this technical report, we present the training framework behind the Yuanbao search agent, designed to achieve search specialization without sacrificing general intelligence. Built upon the Hunyuan3 architecture, our framework combines agentic reinforcement learning for autonomous search with a cross-domain expert On-Policy Distillation (OPD) pipeline. Experts specializing in complementary general-purpose domains are distilled into the search-specialized student, restoring and further enhancing its broad capabilities. Rather than treating specialization and general capability as competing objectives, our hybrid training strategy jointly optimizes both, effectively mitigating the alignment tax. Extensive experiments demonstrate that the resulting model achieves competitive search performance while consistently improving its general-purpose capabilities, providing a favorable balance between specialized execution and broad generalization in real-world search scenarios.

搜索代理强化学习知识蒸馏通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。