arXiv:2512.23717cs.CLcs.AI2025-12被引 1

用多智能体辩论生成隐蔽有害问题,提升大模型安全训练数据质量

HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate

  • 多智能体辩论迭代优化,将明显有害问题转为隐蔽形式
  • 实验显示生成效果显著优于传统基线方法
  • 适合研究模型安全对齐与对抗样本生成的学者

大型语言模型配备安全机制以检测并屏蔽有害查询,但现有对齐方法主要关注显性危险内容,忽视了更隐蔽的威胁。用户常通过隐晦改写保留恶意意图却伪装成无害,导致现有安全训练数据存在显著空白。为此,我们提出HarmTransform,一种多智能体辩论框架,可系统性地将有害查询转化为更具隐蔽性的形式,同时保持其原始恶意目的。该框架通过多个智能体间的迭代批判与优化,生成高质量、隐蔽性强的有害查询变体,可用于改进未来的模型安全对齐训练。实验表明,HarmTransform在生成有效查询变换方面显著优于标准基线。同时分析发现,辩论虽能增强变换效果和隐蔽性,也可能引发话题偏移和不必要的复杂性。这些发现揭示了多智能体辩论在生成全面安全训练数据方面的潜力与局限。

原文摘要 · Abstract (English)

Large language models (LLMs) are equipped with safety mechanisms to detect and block harmful queries, yet current alignment approaches primarily focus on overtly dangerous content and overlook more subtle threats. However, users can often disguise harmful intent through covert rephrasing that preserves malicious objectives while appearing benign, which creates a significant gap in existing safety training data. To address this limitation, we introduce HarmTransform, a multi-agent debate framework for systematically transforming harmful queries into stealthier forms while preserving their underlying harmful intent. Our framework leverages iterative critique and refinement among multiple agents to generate high-quality, covert harmful query transformations that can be used to improve future LLM safety alignment. Experiments demonstrate that HarmTransform significantly outperforms standard baselines in producing effective query transformations. At the same time, our analysis reveals that debate acts as a double-edged sword: while it can sharpen transformations and improve stealth, it may also introduce topic shifts and unnecessary complexity. These insights highlight both the promise and the limitations of multi-agent debate for generating comprehensive safety training data.

模型安全多智能体对抗生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。