arXiv:2608.19579cs.AImath.DS2026-08

用动态系统分析提示与回复的演化,自动识别大模型输出的安全风险。

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

论文配图:Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics
图 1 · 摘自论文原文
  • 基于柯普曼算子构建安全与不安全状态的预测模型
  • 结合提示与回复嵌入,差分残差得分提升检测准确率
  • 适用于黑盒场景,尤其对因果解码器效果更优

大语言模型在高风险应用中日益普及,但其生成有害或违规内容的风险仍难消除。本文将近期用于幻觉检测的动力系统框架拓展至模型安全分类。通过将提示与回复投影至高维嵌入空间,分别为安全与不安全状态拟合柯普曼基预测模型,利用新提出的差分残差分数对比两类模型的预测误差以判断输出安全性。关键贡献在于融合提示与回复嵌入动态,得到能捕捉关键交互模式的柯普曼算子。我们在三个安全基准上使用三种嵌入模型评估该方法。结果表明,引入提示嵌入可稳定提升性能,尤其在关联性违规场景下(如搭配因果解码器的Llama-3),而仅依赖回复的违规则更受益于密集语义嵌入表示。研究为用动力系统分析AI系统提供了新路径,突破了当前以AI建模动力系统的主流范式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.

大模型安全动态系统黑盒检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。