用图像编辑生成动态危险场景,评估视觉语言模型安全决策能力
Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs
- 用图像编辑+布局算法生成包含移动、侵入物体的复杂场景
- 构建含7254张图的MovSafeBench,发现异常物体使模型性能下降显著
- 适合研究自动驾驶中VLM安全性与鲁棒性的研究人员使用
视觉语言模型(VLMs)在自动驾驶和移动系统中应用日益广泛,评估其在复杂环境中的安全决策能力至关重要。然而现有基准未能充分覆盖具有时空动态特性的多样化危险情境,尤其是异常场景。尽管图像编辑模型可用来合成此类风险,但生成结构合理、包含移动、侵入及远距离物体的真实感场景仍具挑战。为此,我们提出HazardForge——一个可扩展的流水线,利用图像编辑模型结合布局决策算法与验证模块生成此类场景。基于HazardForge,我们构建了MovSafeBench,一个多项选择题(MCQ)基准,涵盖13类物体,共7,254张图像与对应问答对,覆盖正常与异常物体。实验表明,在存在异常物体的情境下,VLM性能显著下降,尤其在需要细微运动理解的任务中降幅最大。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are increasingly deployed in autonomous vehicles and mobile systems, making it crucial to evaluate their ability to support safer decision-making in complex environments. However, existing benchmarks inadequately cover diverse hazardous situations, especially anomalous scenarios with spatio-temporal dynamics. While image editing models are a promising means to synthesize such hazards, it remains challenging to generate well-formulated scenarios that include moving, intrusive, and distant objects frequently observed in the real world. To address this gap, we introduce \textbf{HazardForge}, a scalable pipeline that leverages image editing models to generate these scenarios with layout decision algorithms, and validation modules. Using HazardForge, we construct \textbf{MovSafeBench}, a multiple-choice question (MCQ) benchmark comprising 7,254 images and corresponding QA pairs across 13 object categories, covering both normal and anomalous objects. Experiments using MovSafeBench show that VLM performance degrades notably under conditions including anomalous objects, with the largest drop in scenarios requiring nuanced motion understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。