发现语言模型物理安全风险与文本风险可分离,提出新方法精准识别物理越狱。
When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space

- 通过隐藏状态分析揭示物理越狱与文本越狱信号独立
- PRISM模型在安全基准上达87.7%准确率,误报率仅13.7%
- 无需依赖文字模板,适合评估智能体物理安全
大型语言模型(LLMs)越来越多地作为具身智能体的高层规划器,语言上看似无害的指令在物理世界落地后可能变得危险。我们研究这种物理化越狱是否与传统文本越狱属于同一类安全问题。通过对隐藏状态方向分析和随机分割空检验,发现在Qwen2.5-3B/7B/14B/32B、Phi-3.5和SmolLM2中,文本越狱(TJ)与物理越狱(PJ)在模型表示中形成可分离的信号。基于此,我们提出PRISM:一种基于全隐藏状态的单层L2正则逻辑探测器。PRISM在SafeAgentBench上达到86.2–87.7%准确率,误报率(FPR)为11.7–13.7%,而同规模语言模型判断则在24.7–39.0%之间。为验证结果是否受词汇捷径影响,我们引入交互平衡版PhysicalJailbreakBench-2K(PJB-2K),固定2,000行对比集。在底层10,000行池中,词频TFIDF与嵌入层表现仅为随机水平(AUC 0.497和0.500);第25层经独立同分布筛选后,细胞分组交叉验证显示PRISM AUC达0.718,远高于无物理信息标签控制的0.398。在相同2,000行数据上,PRISM预测得0.671平衡准确率,而Qwen2.5从3B到72B模型仅得0.538–0.577,且误报率高。结果支持隐藏状态探测作为超越文本过滤的物理安全表征级方法。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the same safety problem as ordinary textual jailbreak. Through hidden-state direction analysis and random-split null tests, we show that textual jailbreak (TJ) and physical jailbreak (PJ) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on this separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% false-positive rates (FPRs), while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. To test whether the result survives lexical-shortcut controls, we introduce an interaction-balanced revision of PhysicalJailbreakBench-2K (PJB-2K): a fixed 2{,}000-row comparison set sampled by label and physical mechanism from a larger object--site construction. On the underlying 10{,}000-row pool, word-TFIDF and the embedding layer remain at chance (AUC 0.497 and 0.500). At layer 25, selected by an i.i.d. sweep, cell-grouped cross-validation gives PRISM 0.718 AUC, compared with 0.398 for a physics-free label control under the same protocol. On the identical 2{,}000 comparison rows, these PRISM predictions obtain 0.671 balanced accuracy, while Qwen2.5 judges from 3B to 72B obtain 0.538--0.577 and exhibit high FPR. These results support hidden-state probing as a representation-level method for physical safety beyond text moderation, without relying on the near-perfect scores of shortcut-prone paired templates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。