arXiv:2510.21842cs.CVcs.CR2025-10中稿 · ICLR

统一多模态模型记得住图像却说不清,存在记忆与表达的割裂。

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

  • 模型能精准复现电影海报,但描述时关键细节出错。
  • 多种架构在合成数据上均出现此现象,非训练偶然。
  • 提示安全防护只限文本,可能被图像绕过,危及AI安全。

我们提出‘模态失语’现象:当前统一多模态模型虽能准确视觉记忆概念,却无法用文字正确表达,尽管它们同时在图像和文本上训练。实验显示,前沿模型可生成近乎完美的经典电影艺术作品复制品,但在要求提供文本描述时却混淆关键细节。我们在多种架构的合成数据集上通过受控实验验证了该现象的稳定性,确认其为当前统一多模态模型的根本属性,而非训练缺陷。实际中,模态失语可能引入AI安全漏洞——若仅对文本施加安全约束,有害概念仍可通过图像模态传播。我们通过实证表明,仅以文本对齐的模型仍可生成不安全图像。

原文摘要 · Abstract (English)

We present modal aphasia, a systematic dissociation in which current unified multimodal models accurately memorize concepts visually but fail to articulate them in writing, despite being trained on images and text simultaneously. For one, we show that leading frontier models can generate near-perfect reproductions of iconic movie artwork, but confuse crucial details when asked for textual descriptions. We corroborate those findings through controlled experiments on synthetic datasets in multiple architectures. Our experiments confirm that modal aphasia reliably emerges as a fundamental property of current unified multimodal models, not just as a training artifact. In practice, modal aphasia can introduce vulnerabilities in AI safety frameworks, as safeguards applied to one modality may leave harmful concepts accessible in other modalities. We demonstrate this risk by showing how a model aligned solely on text remains capable of generating unsafe images.

多模态模型安全图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。