为深度研究设计四阶段安全防护框架,提升报告安全性与质量。
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
- 引入四阶段防护机制,结合开放域评估确保每阶段安全
- 在GPT-4o等模型上提升16.53%防御成功率,过拒率降至6%
- 提出DRSafeBench基准,支持分阶段安全评测与内容净化
深度研究框架在整合网络信息生成综合报告方面展现出巨大潜力。然而,现有框架缺乏充分的评估流程和阶段化保护机制,通常仅以问答准确率作为评价标准,忽视了报告的可信度、连贯性、广度、深度及安全性。这可能导致有害或恶意来源被纳入最终报告。为此,我们提出DeepResearchGuard框架,包含四阶段安全保障与开放域评估,并构建DRSafeBench这一新型分阶段安全基准。实验表明,在GPT-4o、o4-mini、Gemini-2.5-flash、DeepSeek-v3、GPT-5等模型上,该框架使防御成功率提升16.53%,同时将过拒率降低至6%。通过大量实验验证,DRSafeBench可实现全面的开放域评估与阶段感知防御,有效阻止有害内容传播,系统性提升报告质量且不过度拒绝合理请求。
原文摘要 · Abstract (English)
Deep research frameworks have shown promising capabilities in synthesizing comprehensive reports from web sources. While deep research possesses significant potential to address complex issues through planning and research cycles, existing frameworks are deficient in sufficient evaluation procedures and stage-specific protections. They typically treat evaluation as exact match accuracy of question-answering, but overlook crucial aspects of report quality such as credibility, coherence, breadth, depth, and safety. This oversight may result in hazardous or malicious sources being integrated into the final report. To address this, we introduce DeepResearchGuard, a framework featuring four-stage safeguards with open-domain evaluation, and DRSafeBench, a novel stage-wise safety benchmark. Evaluating across GPT-4o, o4-mini, Gemini-2.5-flash, DeepSeek-v3, GPT-5, DeepResearchGuard improves defense success rates by 16.53% while reducing over-refusal to 6%. Through extensive experiments, we show that DRSafeBench enables comprehensive open-domain evaluation and stage-aware defenses that effectively block harmful content propagation, while systematically improving report quality without excessive over-refusal rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。