arXiv:2606.02423cs.CLcs.LG2026-06

提出新基准与监控框架,防范大模型对话中危害的持续放大。

Investigating and Alleviating Harm Amplification in LLM Interactions

论文配图:Investigating and Alleviating Harm Amplification in LLM Interactions
图 1 · 摘自论文原文
  • 构建多轮危害放大评估基准HarmAmp,覆盖12类真实威胁场景。
  • 提出TrajSafe监控系统,可提前识别并干预潜在有害对话轨迹。
  • 在不降低模型能力的前提下,显著减少多轮对话中的危害输出。

大型语言模型(LLMs)既能作为助手,也可能成为危害放大器,使恶意用户通过多轮交互实现超出自身能力的有害目标。这一风险体现在两个维度:一是让新手获得专业领域危害内容的创作能力,二是将有害操作规模化至人力无法企及的程度。现有研究常忽视模型在多轮对话中对危害的累积效应。本文提出HarmAmp,一个涵盖十二类风险的多轮危害放大评估基准,每个场景均基于真实威胁,满足实质性放大、操作特异性与多轮必要性要求。进一步提出TrajSafe,一种主动监控机制,通过探查用户真实意图并引导模型走向安全生成,提前干预有害路径。大量实验表明,TrajSafe显著降低了多轮交互中的危害程度,同时保持低误拒率和原模型的通用能力。本工作为缓解大模型交互中的复杂安全风险提供了可行范式。

原文摘要 · Abstract (English)

Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve harmful outcomes beyond their capabilities through extended interactions. This risk manifests along two axes, i.e., democratizing domain expertise that allows novices to produce specialized harmful content, and scaling harmful operations at volumes that manual effort cannot match. Existing works, however, often overlook how LLMs compound harm across multi-turn conversations. We introduce HarmAmp, a new benchmark for multi-turn harm amplification scenarios spanning twelve risk categories. Each scenario is grounded in real-world threats and satisfies rigorous criteria, i.e., substantive amplification, operational specificity, and multi-turn necessity. We further propose TrajSafe, a proactive monitor that anticipates harmful trajectories and intervenes through actions such as probing users' genuine intents and steering the models towards safer completion. Our extensive experiments demonstrate that TrajSafe significantly reduces the harmfulness incurred in multi-turn interactions while preserving a low over-refusal rate and the target model's general capabilities. Our work offers a promising paradigm to alleviate the nuanced safety risks in LLM interactions.

大模型安全多轮对话危害放大主动防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。