arXiv:2603.27332cs.CV2026-03被引 2

统一多模态模型因生成与理解互耦,易被攻击放大安全风险。

Unsafe by Reciprocity: How Generation-Understanding Coupling Undermines Safety in Unified Multimodal Models

  • 提出新攻击框架RICE,利用生成与理解双向交互漏洞
  • 实验显示双向攻击成功率高,安全风险显著提升
  • 揭示统一模型固有安全隐患,适合安全研究者关注

大型语言模型(LLMs)和文本到图像(T2I)模型的发展催生了统一多模态模型(UMMs),其理解与生成功能在共享架构中紧密耦合。现有研究认为这种互惠机制通过共享表征和联合优化提升了跨功能性能,但其安全影响仍不明确,因多数安全研究孤立分析理解与生成功能。本文探究这种跨功能互惠是否构成结构性安全漏洞。我们提出RICE:基于互惠交互的跨功能攻击框架,系统评估生成→理解(G-U)与理解→生成(U-G)攻击路径,发现不安全中间信号可在模态间传播并放大风险。大量实验表明双向攻击成功率高,揭示了UMMs中此前未被注意的安全弱点。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) and Text-to-Image (T2I) models have led to the emergence of Unified Multimodal Models (UMMs), where multimodal understanding and image generation are tightly integrated within a shared architecture. Prior studies suggest that such reciprocity enhances cross-functionality performance through shared representations and joint optimization. However, the safety implications of this tight coupling remain largely unexplored, as existing safety research predominantly analyzes understanding and generation functionalities in isolation. In this work, we investigate whether cross-functionality reciprocity itself constitutes a structural source of vulnerability in UMMs. We propose RICE: Reciprocal Interaction-based Cross-functionality Exploitation, a novel attack paradigm that explicitly exploits bidirectional interactions between understanding and generation. Using this framework, we systematically evaluate Generation-to-Understanding (G-U) and Understanding-to-Generation (U-G) attack pathways, demonstrating that unsafe intermediate signals can propagate across modalities and amplify safety risks. Extensive experiments show high Attack Success Rates (ASR) in both directions, revealing previously overlooked safety weaknesses inherent to UMMs.

多模态安全攻击方法生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。