arXiv:2603.29759cs.CVcs.AI2026-03

构建真实场景下的安全风险评估基准,推动视觉语言模型更可靠地识别居家隐患。

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

  • 融合真实数据与生成内容,构建涵盖6.5万条问答的训练集和1707条挑战性测试集。
  • 22个主流模型在新基准上平均性能提升18.3分,验证了其评估价值。
  • 适合关注智能安全、多模态推理与模型鲁棒性的研究者使用。

视觉语言模型(VLMs)在室内安全风险评估中的应用日益广泛,但现有基准存在三大局限:(1)过度依赖仿真软件生成的合成数据,与真实环境存在显著领域差距;(2)安全任务过于简化,对风险类型和场景设置人为限制,削弱模型泛化能力;(3)缺乏严谨的评估协议,难以全面衡量模型在复杂家庭场景中的表现。为此,我们提出TSHA(可信安全风险评估基准),包含66,668条经验证的问答对,其中64,961条来自现有室内数据集、互联网图像、AIGC图像、新采集图像及鸿蒙全景图像。测试集含1,707条高质量问答,不仅包含训练分布的精选子集,还引入Sora生成视频和鸿蒙全景图像中的多风险场景,用于评估模型在复杂情境下的鲁棒性。对22个主流VLMs的实验表明,当前模型在安全评估任务中普遍缺乏稳健能力。值得注意的是,基于TSHA训练的模型在测试集上性能最高提升18.3分,并在其他基准上展现出更强泛化性,凸显该基准的重大贡献与价值。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental limitations: (1) heavy reliance on synthetic datasets constructed via simulation software, creating a significant domain gap with real-world environments; (2) oversimplified safety tasks with artificial constraints on hazard and scene types, thereby limiting model generalization; and (3) absence of rigorous evaluation protocols to thoroughly assess model capabilities in complex home safety scenarios. To address these challenges, we introduce TSHA (\textbf{T}rustworthy \textbf{S}afety \textbf{H}azards \textbf{A}ssessment), a comprehensive benchmark comprising 66,668 validated question-answer pairs, including 64,961 carefully curated training QA pairs drawn from existing indoor datasets, internet frames/images, AIGC images, newly captured images, and Hunyuan panoramic images. This benchmark also includes a highly challenging test set with 1,707 QA pairs, comprising not only a carefully selected subset from the training distribution but also newly added Sora-generated videos and Hunyuan panoramic images containing multiple safety hazards, used to evaluate the model's robustness in complex safety scenarios. Extensive experiments on 22 popular VLMs demonstrate that current VLMs lack robust capabilities for safety hazard assessment. Importantly, models trained on the TSHA training set achieve a significant performance improvement of up to +18.3 points on the TSHA test set and also exhibit enhanced generalizability across other benchmarks, underscoring the substantial contribution and importance of the TSHA benchmark.

视觉语言模型安全评估基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。