arXiv:2601.04043cs.CL2026-01ACL被引 3

构建真实场景多模态安全评测集,揭示大模型在日常安全判断中的严重缺陷。

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

  • 设计包含2013个真实图像-文本对的多模态安全基准数据集
  • 顶尖模型在危险问题上仅57.2%给出安全响应,多数仍会错误协助
  • 强调跨模态推理与真实风险感知,适合评估模型真实安全能力

随着多模态大语言模型(MLLMs)成为人类生活不可或缺的助手,其生成的不安全内容对人类行为构成持续威胁。为研究和评估MLLMs对人类日常行为的安全影响,我们提出SaLAD——一个包含2,013个真实世界图像-文本样本的多模态安全基准,覆盖10类常见场景,兼顾危险情境与过度敏感案例。该数据集强调真实风险暴露、真实视觉输入与细粒度跨模态推理,确保安全风险无法仅通过文本推断。我们还提出基于安全警告的评估框架,鼓励模型提供明确、有信息量的安全提示,而非泛化拒绝。对18个MLLMs的测试表明,表现最好的模型在危险查询上的安全响应率仅为57.2%。即使主流安全对齐方法也未能有效提升模型在此场景下的表现,暴露出当前MLLMs在识别日常危险行为方面的显著脆弱性。数据集已公开于https://github.com/xinyuelou/SaLAD。

原文摘要 · Abstract (English)

As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to human behavior, perpetually overhanging human society like a sword of Damocles. To investigate and evaluate the safety impact of MLLMs responses on human behavior in daily life, we introduce SaLAD, a multimodal safety benchmark which contains 2,013 real-world image-text samples across 10 common categories, with a balanced design covering both unsafe scenarios and cases of oversensitivity. It emphasizes realistic risk exposure, authentic visual inputs, and fine-grained cross-modal reasoning, ensuring that safety risks cannot be inferred from text alone. We further propose a safety-warning-based evaluation framework that encourages models to provide clear and informative safety warnings, rather than generic refusals. Results on 18 MLLMs demonstrate that the top-performing models achieve a safe response rate of only 57.2% on unsafe queries. Moreover, even popular safety alignment methods limit effectiveness of the models in our scenario, revealing the vulnerabilities of current MLLMs in identifying dangerous behaviors in daily life. Our dataset is available at https://github.com/xinyuelou/SaLAD.

多模态安全大模型评测风险识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。