评估生成智能体在多模态环境中的安全表现,揭示其信任视觉信息的倾向。
Multimodal Safety Evaluation in Generative Agent Social Simulations
- 构建可复现的仿真框架,从安全、检测、社交三维度量化智能体行为。
- 多模态矛盾下仅55%计划能正确修正,视觉误导使45%不安全行为被接受。
- 适合研究多模态安全、智能体社会性与可信推理的研究者参考。
生成智能体能否在多模态环境中被信赖?尽管大语言模型和视觉-语言模型使智能体能在复杂场景中自主行动,但其在跨模态安全、一致性与信任推理方面仍存在局限。本文提出一个可复现的仿真框架,从三个维度评估智能体:(1) 安全性随时间提升,包括文本-视觉场景中的迭代计划修订;(2) 多类社交情境下不安全行为的检测;(3) 社交动态,以交互次数和社交交换接受率衡量。智能体配备分层记忆、动态规划与多模态感知,并集成SocialMetrics工具包,量化计划修订、不安全到安全的转换及信息扩散。实验显示,智能体虽可识别直接多模态矛盾,却常无法将局部修正与全局安全对齐,仅55%计划修正成功。八次仿真中,Claude、GPT-4o mini、Qwen-VL三模型平均不安全转安全率分别为75%、55%、58%。整体性能从多风险场景中GPT-4o mini的20%到局部场景如火灾/高温中Claude的98%不等。值得注意的是,45%的不安全行为在搭配误导性图像时被接受,表明智能体对图像存在强烈过度信任。这些发现暴露了当前架构的关键缺陷,并提供了一个可复现的平台,用于研究多模态安全、一致性和社交动态。
原文摘要 · Abstract (English)
Can generative agents be trusted in multimodal environments? Despite advances in large language and vision-language models that enable agents to act autonomously and pursue goals in rich settings, their ability to reason about safety, coherence, and trust across modalities remains limited. We introduce a reproducible simulation framework for evaluating agents along three dimensions: (1) safety improvement over time, including iterative plan revisions in text-visual scenarios; (2) detection of unsafe activities across multiple categories of social situations; and (3) social dynamics, measured as interaction counts and acceptance ratios of social exchanges. Agents are equipped with layered memory, dynamic planning, multimodal perception, and are instrumented with SocialMetrics, a suite of behavioral and structural metrics that quantifies plan revisions, unsafe-to-safe conversions, and information diffusion across networks. Experiments show that while agents can detect direct multimodal contradictions, they often fail to align local revisions with global safety, reaching only a 55 percent success rate in correcting unsafe plans. Across eight simulation runs with three models - Claude, GPT-4o mini, and Qwen-VL - five agents achieved average unsafe-to-safe conversion rates of 75, 55, and 58 percent, respectively. Overall performance ranged from 20 percent in multi-risk scenarios with GPT-4o mini to 98 percent in localized contexts such as fire/heat with Claude. Notably, 45 percent of unsafe actions were accepted when paired with misleading visuals, showing a strong tendency to overtrust images. These findings expose critical limitations in current architectures and provide a reproducible platform for studying multimodal safety, coherence, and social dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。