用自动合成数据训练智能体,让大模型更会应对检索失败和干扰。
RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis
- 通过构建带干扰的信息树,模拟真实检索环境的复杂性。
- 训练出的模型在噪声环境下表现显著优于现有基线。
- 适合研究智能体推理、检索增强生成的开发者与学者。
智能体检索增强生成(Agentic RAG)使大语言模型能够自主规划并检索信息以解决复杂问题。然而,由于高质量训练数据稀缺,尤其难以反映真实检索环境中的噪声与复杂性,模型发展受限。传统人工标注成本高且无法捕捉动态推理策略。为此,我们提出RAGShaper,一种自动化数据合成框架,用于构建复杂的RAG任务与鲁棒智能体轨迹。RAGShaper引入InfoCurator构建富含对抗性干扰的信息树,覆盖感知与认知层级;同时设计约束导航策略,迫使教师智能体面对干扰,从而显式生成错误修正与噪声排除的轨迹。大量实验表明,基于合成语料训练的模型在高噪声和复杂检索任务中显著优于现有基线,展现出更强的鲁棒性。
原文摘要 · Abstract (English)
Agentic Retrieval-Augmented Generation (RAG) empowers large language models to autonomously plan and retrieve information for complex problem-solving. However, the development of robust agents is hindered by the scarcity of high-quality training data that reflects the noise and complexity of real-world retrieval environments. Conventional manual annotation is unscalable and often fails to capture the dynamic reasoning strategies required to handle retrieval failures. To bridge this gap, we introduce RAGShaper, a novel data synthesis framework designed to automate the construction of RAG tasks and robust agent trajectories. RAGShaper incorporates an InfoCurator to build dense information trees enriched with adversarial distractors spanning Perception and Cognition levels. Furthermore, we propose a constrained navigation strategy that forces a teacher agent to confront these distractors, thereby eliciting trajectories that explicitly demonstrate error correction and noise rejection. Comprehensive experiments confirm that models trained on our synthesized corpus significantly outperform existing baselines, exhibiting superior robustness in noise-intensive and complex retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。