arXiv:2604.14865cs.CLcs.CR2026-04被引 4

通过多段证据聚合提升大模型有害意图探测的鲁棒性

Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

论文配图:Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
图 1 · 摘自论文原文
  • 用多个一致证据令牌替代单一高分词,避免误报
  • 在1%误报率下真阳性率提升35.55%,AUROC达98.85%
  • 适用于对抗性加密攻击,可直接复用基础模型探针

大型语言模型在高风险的化学、生物、辐射和核(CBRN)领域面临持续演化的越狱攻击。尽管流式探测能实现实时监控,但仍存在系统性错误。我们发现核心问题在于:现有方法依赖少数高分令牌,导致敏感词出现在良性语境中时产生误报。为此,提出一种流式探测目标,要求多个证据令牌持续支持同一预测,而非依赖孤立的分数突增。该机制促使检测基于信号聚合而非单令牌线索。在固定1%误报率下,本方法相较强基线提升35.55%相对真阳性率;即使基线已接近饱和(AUROC=97.40%),仍实现显著增益。进一步发现,探测注意力或MLP激活层优于残差流特征。即便面对新型字符级加密的对抗微调攻击,原生模型训练的探测器仍可“即插即用”,实现超过98.85%的AUROC。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly exposed to adaptive jailbreaking, particularly in high-stakes Chemical, Biological, Radiological, and Nuclear (CBRN) domains. Although streaming probes enable real-time monitoring, they still make systematic errors. We identify a core issue: existing methods often rely on a few high-scoring tokens, leading to false alarms when sensitive CBRN terms appear in benign contexts. To address this, we introduce a streaming probing objective that requires multiple evidence tokens to consistently support a prediction, rather than relying on isolated spikes. This encourages more robust detection based on aggregated signals instead of single-token cues. At a fixed 1% false-positive rate, our method improves the true-positive rate by 35.55% relative to strong streaming baselines. We further observe substantial gains in AUROC, even when starting from near-saturated baseline performance (AUROC = 97.40%). We also show that probing Attention or MLP activations consistently outperforms residual-stream features. Finally, even when adversarial fine-tuning enables novel character-level ciphers, harmful intent remains detectable: probes developed for the base LLMs can be applied ``plug-and-play'' to these obfuscated attacks, achieving an AUROC of over 98.85%.

大模型安全意图探测流式监控鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。