通过分层测试用例生成与多轮交互,全面发现大模型潜在风险。
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
- 基于细粒度风险分类的自上而下生成方法,提升测试多样性。
- 结合微调与强化学习,实现类人多轮对抗探测。
- 为模型对齐提供系统性漏洞洞察,适合安全评估人员使用。
自动化红队测试是识别大语言模型(LLMs)偏移行为的有效方法。现有方法多关注攻击成功率,忽视测试用例的全面覆盖;且多数仅限单轮测试,无法捕捉真实人机交互中的多轮动态。为此,我们提出HARM(Holistic Automated Red teaMing),采用可扩展的细粒度风险分类体系,通过自上而下的方式生成多样化测试用例。同时,引入新颖的微调策略与强化学习技术,实现类人化的多轮对抗探测。实验表明,该框架能更系统地揭示模型漏洞,并为对齐过程提供更精准的指导。
原文摘要 · Abstract (English)
Automated red teaming is an effective method for identifying misaligned behaviors in large language models (LLMs). Existing approaches, however, often focus primarily on improving attack success rates while overlooking the need for comprehensive test case coverage. Additionally, most of these methods are limited to single-turn red teaming, failing to capture the multi-turn dynamics of real-world human-machine interactions. To overcome these limitations, we propose HARM (Holistic Automated Red teaMing), which scales up the diversity of test cases using a top-down approach based on an extensible, fine-grained risk taxonomy. Our method also leverages a novel fine-tuning strategy and reinforcement learning techniques to facilitate multi-turn adversarial probing in a human-like manner. Experimental results demonstrate that our framework enables a more systematic understanding of model vulnerabilities and offers more targeted guidance for the alignment process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。