用多角色协作动态评估大模型风险,识别更准更全面。
RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
- 分三类风险空间,让不同角色专注识别显性、隐性与非风险内容。
- 在800个挑战样本上,风险识别准确率比最强基线高28.87%。
- 适合关注模型安全评估的工程师与研究人员使用。
现有大语言模型安全评估方法存在评价者偏见和因模型同质化导致的漏检问题,削弱了风险评估的鲁棒性。本文重新审视风险评估范式,提出一个理论框架,将潜在风险概念空间分解为三个互斥子空间:显性风险子空间(直接违反安全规范)、隐性风险子空间(需上下文推理才能识别的潜在恶意内容)和非风险子空间。进一步提出RADAR框架,通过四个专业化互补角色的多轮辩论机制与动态更新策略,实现风险概念分布的自我演化。该方法兼顾显性和隐性风险覆盖,有效缓解评价者偏见。为验证有效性,构建包含800个挑战案例的评估数据集。在该测试集及公开基准上的实验表明,RADAR在准确性、稳定性与自评风险敏感度等维度均显著优于基线方法,尤其在风险识别准确率上相较最强基线提升28.87%。
原文摘要 · Abstract (English)
Existing safety evaluation methods for large language models (LLMs) suffer from inherent limitations, including evaluator bias and detection failures arising from model homogeneity, which collectively undermine the robustness of risk evaluation processes. This paper seeks to re-examine the risk evaluation paradigm by introducing a theoretical framework that reconstructs the underlying risk concept space. Specifically, we decompose the latent risk concept space into three mutually exclusive subspaces: the explicit risk subspace (encompassing direct violations of safety guidelines), the implicit risk subspace (capturing potential malicious content that requires contextual reasoning for identification), and the non-risk subspace. Furthermore, we propose RADAR, a multi-agent collaborative evaluation framework that leverages multi-round debate mechanisms through four specialized complementary roles and employs dynamic update mechanisms to achieve self-evolution of risk concept distributions. This approach enables comprehensive coverage of both explicit and implicit risks while mitigating evaluator bias. To validate the effectiveness of our framework, we construct an evaluation dataset comprising 800 challenging cases. Extensive experiments on our challenging testset and public benchmarks demonstrate that RADAR significantly outperforms baseline evaluation methods across multiple dimensions, including accuracy, stability, and self-evaluation risk sensitivity. Notably, RADAR achieves a 28.87% improvement in risk identification accuracy compared to the strongest baseline evaluation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。