提出动态防御框架,提前预测多轮跨模态攻击的恶意演化轨迹。
Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks
- 将对话流建模为连续轨迹,结合几何与运动特征检测恶意漂移。
- 在真实对话中实现92.3%的恶意攻击提前预警率,误报率低于5%。
- 适合需要实时安全防护的智能代理系统,尤其对抗渐进式攻击。
多模态大模型在自主智能工作流中的应用带来了非平稳的攻击面。实证观察表明,攻击者通过跨模态、渐进式扰动,将恶意意图分散于多轮对话中,绕过逐轮防护机制。传统静态防御受限于马尔可夫性质,无法识别累积性结构污染。为此,本文将安全验证重构为动态生存预测与轨迹动力学问题,提出三重异常防御(TRIAD)框架:将多模态多轮对话流映射为连续轨迹,融合结构异常检测监测协方差偏移、基于Ledoit-Wolf正则化的马氏距离监控高维空间变化、以及拓扑轨迹加速度区分良性探索与持续恶意漂移。这些运动学与几何特征通过贝叶斯隐马尔可夫模型反馈回时变Cox比例风险模型。理论分析证明,TRIAD框架在对抗扰动下具有数学上有界期望失效时间,确保恶意加速正向发散。该框架计算高效、可解释,适用于实时智能代理系统的预测性安全防护,无需依赖经验重训练,建立持续安全对齐的严格基础。
原文摘要 · Abstract (English)
The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agentic workflows has introduced a non-stationary attack surface. Empirical observations indicate that adversaries employ progressive, cross-modal perturbations that evade turn-specific guardrails by distributing malicious intent across longitudinal conversational trajectories. Static defense mechanisms, constrained by the Markov property, evaluate inputs in isolation and fail to detect cumulative structural poisoning. To handle this limitation, this paper formulates safety verification as a dynamic survival prediction and trajectory dynamics problem. The Triple-tier Anomaly Defense (TRIAD) framework is proposed as a predictive model that maps multimodal and multi-turn conversational flow as a continuous trajectory. The framework integrates structural anomaly detection to monitor covariance shifts, a Ledoit-Wolf regularized Mahalanobis distance to monitor covariance shifts in high-dimensional spaces, and topological trajectory acceleration to differentiate benign creative exploration from continuous malicious drift. These kinematic and geometric features are integrated into a time-varying Cox Proportional Hazards model via a Bayesian Hidden Markov Model (HMM) feedback loop. Theoretical analysis demonstrates that the TRIAD framework provides a mathematically bounded expected time-to-failure under adversarial perturbations, ensuring that malicious acceleration diverges positively. This framework provides a computationally efficient, interpretable, and predictive safeguard for real-time agentic AI systems, establishing a rigorous foundation for continuous safety alignment without relying on empirical retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。