arXiv:2601.06049cs.CYcs.AI2026-01

上传带水印的版权图并要求去除,后续所有图像生成请求都会被拒绝。

The Violation State: Safety State Persistence in a Multimodal Language Model Interface

  • 模型在拒绝版权图去水印后,持续拒绝后续无关图像生成请求。
  • 40次会话中,污染组120次图像生成全被拒(96.67%),对照组0次被拒。
  • 揭示了多模态对话中的安全状态持续现象,适合关注系统安全设计的研究者。

多模态人工智能系统将文本生成、图像生成等功能集成于单一对话界面。这些系统通过安全机制防止违规操作,如移除受版权保护图像的水印。尽管单轮拒绝是预期行为,但安全过滤器与会话级状态之间的交互机制尚不明确。本研究在ChatGPT(GPT-5.1)网页界面中记录到可复现的行为效应。采用人工执行以捕捉生产系统的真实用户端安全行为,而非孤立的API组件。当会话以上传受版权保护的图像并请求去除水印开始,模型正确拒绝后,后续所有图像生成请求(即使与原请求无关)均被拒绝,直至会话结束。值得注意的是,纯文本请求(如生成Python函数)仍可成功。在40次手动会话中(30次污染组,10次对照组),污染组120次图像生成请求全部被拒(96.67%),对照组40次均未被拒(Fisher精确检验p < 0.0001)。所有会话使用相同固定提示顺序,确保序列一致性。我们将其称为‘安全状态持续’:一种因版权拒绝引发的会话级过度泛化现象。本文呈现为行为观察,非架构性主张。讨论可能解释、方法局限(仅一个模型、单一界面)及对多模态可靠性、用户体验和会话级安全系统设计的影响。结果呼吁进一步探究多模态AI系统中会话级安全机制的交互特性。

原文摘要 · Abstract (English)

Multimodal AI systems integrate text generation, image generation, and other capabilities within a single conversational interface. These systems employ safety mechanisms to prevent disallowed actions, including the removal of watermarks from copyrighted images. While single-turn refusals are expected, the interaction between safety filters and conversation-level state is not well understood. This study documents a reproducible behavioral effect in the ChatGPT (GPT-5.1) web interface. Manual execution was chosen to capture the exact user-facing safety behavior of the production system, rather than isolated API components. When a conversation begins with an uploaded copyrighted image and a request to remove a watermark, which the model correctly refuses, subsequent prompts to generate unrelated, benign images are refused for the remainder of the session. Importantly, text-only requests (e.g., generating a Python function) continue to succeed. Across 40 manually run sessions (30 contaminated and 10 controls), contaminated threads showed 116/120 image-generation refusals (96.67%), while control threads showed 0/40 refusals (Fisher's exact p < 0.0001). All sessions used an identical fixed prompt order, ensuring sequence uniformity across conditions. We describe this as safety-state persistence: a form of conversational over-generalization in which a copyright refusal influences subsequent, unrelated image-generation behavior. We present these findings as behavioral observations, not architectural claims. We discuss possible explanations, methodological limitations (single model, single interface), and implications for multimodal reliability, user experience, and the design of session-level safety systems. These results motivate further examination of session-level safety interactions in multimodal AI systems.

多模态安全机制会话状态版权防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。