arXiv:2602.21447cs.CRcs.AI2026-02被引 3

通过隐变量建模对抗意图,实现多阶段推理的安全防护。

Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG

  • 将攻击意图视为隐藏状态,用分层大模型推理更新信任信念。
  • 在43,774次测试中,攻击成功率降低6.5倍,性能损失极小。
  • 适合需要高安全性的多模态智能体系统使用。

当前针对多模态智能体RAG的无状态防御无法检测跨检索、规划和生成阶段分布的恶意语义。本文将该安全挑战建模为部分可观测马尔可夫决策过程(POMDP),其中对抗意图是需从多阶段噪声观测中推断的隐变量。提出MMA-RAG^T,一种由模块化信任代理(MTA)控制的运行时调控框架,通过结构化LLM推理维护近似信念状态。作为与模型无关的叠加层,MMA-RAG^T可配置内部检查点,实现状态化纵深防御。在43,774个实例上的评估显示,相比未受保护基线,攻击成功率平均降低6.50倍,且实用性损耗可忽略。关键因子消融实验验证了理论边界:状态性与空间覆盖均单独必要(分别提升26.4和13.6个百分点),而在检查点检测完全相关时,无状态多点干预的边际收益为零。

原文摘要 · Abstract (English)

Current stateless defences for multimodal agentic RAG fail to detect adversarial strategies that distribute malicious semantics across retrieval, planning, and generation components. We formulate this security challenge as a Partially Observable Markov Decision Process (POMDP), where adversarial intent is a latent variable inferred from noisy multi-stage observations. We introduce MMA-RAG^T, an inference-time control framework governed by a Modular Trust Agent (MTA) that maintains an approximate belief state via structured LLM reasoning. Operating as a model-agnostic overlay, MMA-RAGT mediates a configurable set of internal checkpoints to enforce stateful defence-in-depth. Extensive evaluation on 43,774 instances demonstrates a 6.50x average reduction factor in Attack Success Rate relative to undefended baselines, with negligible utility cost. Crucially, a factorial ablation validates our theoretical bounds: while statefulness and spatial coverage are individually necessary (26.4 pp and 13.6 pp gains respectively), stateless multi-point intervention can yield zero marginal benefit under homogeneous stateless filtering when checkpoint detections are perfectly correlated.

安全防御多模态RAG信任推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。