arXiv:2604.19859cs.LGcs.AI2026-04被引 4

用1万条公开数据训练出40亿参数的边缘科研助手,性能超越90亿参数模型。

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

论文配图:DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data
图 1 · 摘自论文原文
  • 通过两阶段训练提升小模型数据质量与利用效率,构建高效科研代理。
  • 在多个深度研究基准上,40亿参数模型表现优于90亿参数同类模型。
  • 适合关注边缘部署、低资源科研自动化及小模型潜力的研究者。

基于小型语言模型的边缘规模深度研究代理因成本低、延迟小、隐私性好而具有实际部署前景。本文研究如何在有限开放数据下训练强大小型研究代理,以提升数据质量与利用率。我们提出 DR-Venus,一个完全基于开放数据构建的前沿 40 亿参数深度研究代理,适用于边缘部署。训练分两阶段:第一阶段采用代理式监督微调(SFT),结合严格数据清洗与长轨迹重采样,提升数据质量与利用效率;第二阶段应用代理式强化学习(RL),通过基于信息增益和格式感知正则化的回合级奖励设计,增强监督密度与回合内信用分配,从而提升长期任务执行可靠性。整个模型仅依赖约 10,000 条开放数据,显著优于现有 90 亿参数以下的代理模型,同时缩小了与 300 亿参数系统间的差距。进一步分析表明,40 亿参数模型已具备惊人性能潜力,凸显小型模型的部署前景与该场景下测试时扩展的价值。我们开源模型、代码与关键训练配方,支持可复现的边缘规模深度研究代理研究。

原文摘要 · Abstract (English)

Edge-scale deep research agents based on small language models are attractive for real-world deployment due to their advantages in cost, latency, and privacy. In this work, we study how to train a strong small deep research agent under limited open-data by improving both data quality and data utilization. We present DR-Venus, a frontier 4B deep research agent for edge-scale deployment, built entirely on open data. Our training recipe consists of two stages. In the first stage, we use agentic supervised fine-tuning (SFT) to establish basic agentic capability, combining strict data cleaning with resampling of long-horizon trajectories to improve data quality and utilization. In the second stage, we apply agentic reinforcement learning (RL) to further improve execution reliability on long-horizon deep research tasks. To make RL effective for small agents in this setting, we build on IGPO and design turn-level rewards based on information gain and format-aware regularization, thereby enhancing supervision density and turn-level credit assignment. Built entirely on roughly 10K open-data, DR-Venus-4B significantly outperforms prior agentic models under 9B parameters on multiple deep research benchmarks, while also narrowing the gap to much larger 30B-class systems. Our further analysis shows that 4B agents already possess surprisingly strong performance potential, highlighting both the deployment promise of small models and the value of test-time scaling in this setting. We release our models, code, and key recipes to support reproducible research on edge-scale deep research agents.

小模型科研代理边缘计算强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。