测试大模型主动识别风险的能力,发现其表现不稳定。
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
- 构建416个跨模态场景的主动安全评估基准
- 顶尖模型在图像和文本任务中准确率仅71%和64%
- 适合关注AI主动防护能力的研究者与开发者
人类在日常生活中常因安全意识不足而无法及时识别风险。相较于被动响应的系统,主动式安全人工智能能更早预警潜在危险。本文提出主动安全基准(PaSBench),涵盖416个跨模态场景(128个图像序列、288个文本日志),覆盖5个高危领域。对36个先进模型的评估显示,顶级模型Gemini-2.5-pro在图像和文本任务中分别达到71%和64%的准确率,但在重复测试中仍遗漏45%-55%的风险。失败分析表明,模型主要受限于不稳定的主动推理能力,而非知识缺失。本研究建立了主动安全评估基准,提供了系统性证据,并指明了可靠防护型AI的发展方向。数据集已开源:https://huggingface.co/datasets/Youliang/PaSBench。
原文摘要 · Abstract (English)
Human safety awareness gaps often prevent the timely recognition of everyday risks. In solving this problem, a proactive safety artificial intelligence (AI) system would work better than a reactive one. Instead of just reacting to users' questions, it would actively watch people's behavior and their environment to detect potential dangers in advance. Our Proactive Safety Bench (PaSBench) evaluates this capability through 416 multimodal scenarios (128 image sequences, 288 text logs) spanning 5 safety-critical domains. Evaluation of 36 advanced models reveals fundamental limitations: Top performers like Gemini-2.5-pro achieve 71% image and 64% text accuracy, but miss 45-55% risks in repeated trials. Through failure analysis, we identify unstable proactive reasoning rather than knowledge deficits as the primary limitation. This work establishes (1) a proactive safety benchmark, (2) systematic evidence of model limitations, and (3) critical directions for developing reliable protective AI. We believe our dataset and findings can promote the development of safer AI assistants that actively prevent harm rather than merely respond to requests. Our dataset can be found at https://huggingface.co/datasets/Youliang/PaSBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。