arXiv:2504.13201cs.CRcs.LG2025-04被引 4

动态旋转语义子空间,提升智能体系统防御攻击时的生成质量。

CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation

  • 针对任务动态构建安全语义子空间,实现自适应防护。
  • 在相近防御成功率下,显著提升生成内容质量和任务遵循度。
  • 适合注重生成质量与系统可用性的智能体应用。

大型语言模型(LLMs)广泛用于具身智能(EI)系统中的任务理解与动作规划,但其使用大幅增加了遭受越狱攻击的风险。现有推理阶段防御方法依赖对中间表示的静态干预,常导致生成质量下降并影响任务指令遵循,降低系统在具身智能场景下的可用性。本文提出一种动态防御框架:针对每个具身智能推理请求,动态构建特定任务的安全语义子空间,将隐藏状态投影至最相关方向,并通过SLERP旋转实现自适应安全控制。在相近防御成功率下,该方法保持生成质量,提升系统可用性,降低调优成本,并增强在具身智能场景中的鲁棒性。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely used for task understanding and action planning in embodied intelligence (EI) systems, but their adoption substantially increases vulnerability to jailbreak attacks. While recent work explores inference-time defenses, existing methods rely on static interventions on intermediate representations, which often degrade generation quality and impair adherence to task instructions, reducing system usability in EI settings. We propose a dynamic defense framework. For each EI inference request, we dynamically construct a task-specific safety-semantic subspace, project its hidden state to the most relevant direction, and apply SLERP rotation for adaptive safety control. At comparable defense success rates, our method preserves generation quality, improves usability, reduces tuning cost, and strengthens robustness in EI scenarios.

具身智能越狱防御动态防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。