arXiv:2603.16445cs.AI2026-03

视觉干扰让大模型道德判断失控,突破文本安全机制

Visual Distraction Undermines Moral Reasoning in Vision-Language Models

  • 构建多模态道德困境测试框架,分离视觉与语境变量
  • 视觉输入使模型转向直觉反应,绕过文本安全约束
  • 揭示视觉语言模型在道德判断上的关键安全漏洞

道德推理是实现安全人工智能的基础,但随着AI系统从文本助手演变为具身代理,跨模态一致性变得尤为关键。现有安全技术在文本场景中表现良好,但在视觉输入下的泛化能力仍存疑。现有道德评估基准多为纯文本形式,缺乏对影响道德决策变量的系统控制。本文展示,视觉输入会从根本上改变当前顶尖视觉-语言模型(VLMs)的道德判断,使其绕过基于文本的安全机制。我们提出基于道德基础理论(MFT)的道德困境模拟(MDS)框架,通过正交操控视觉与语境变量,实现机制性分析。评估结果表明,视觉模态激活了类直觉路径,取代了文本场景中更审慎、更安全的推理模式。这一发现暴露了语言调优的安全过滤器在视觉处理中失效的关键脆弱性,凸显了多模态安全对齐的紧迫需求。

原文摘要 · Abstract (English)

Moral reasoning is fundamental to safe Artificial Intelligence (AI), yet ensuring its consistency across modalities becomes critical as AI systems evolve from text-based assistants to embodied agents. Current safety techniques demonstrate success in textual contexts, but concerns remain about generalization to visual inputs. Existing moral evaluation benchmarks rely on textonly formats and lack systematic control over variables that influence moral decision-making. Here we show that visual inputs fundamentally alter moral decision-making in state-of-the-art (SOTA) Vision-Language Models (VLMs), bypassing text-based safety mechanisms. We introduce Moral Dilemma Simulation (MDS), a multimodal benchmark grounded in Moral Foundation Theory (MFT) that enables mechanistic analysis through orthogonal manipulation of visual and contextual variables. The evaluation reveals that the vision modality activates intuition-like pathways that override the more deliberate and safer reasoning patterns observed in text-only contexts. These findings expose critical fragilities where language-tuned safety filters fail to constrain visual processing, demonstrating the urgent need for multimodal safety alignment.

道德推理视觉语言模型安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。