arXiv:2603.11975cs.CVcs.AI2026-03被引 3

评测视觉语言模型在家庭场景中识别危险行为的能力,提出实时安全防护架构。

HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios

  • 构建混合仿真与视频生成的动态基准,覆盖438个家庭场景案例。
  • 提出分层双脑架构,实现低延迟与高精度的实时安全监控。
  • 揭示现有视觉语言模型在动态危险检测中的关键瓶颈。

embodied agents 的快速发展推动了家用机器人在真实环境中的部署。然而,与结构化的工业环境不同,家庭空间存在不可预测的安全风险,系统感知延迟和常识知识缺失可能导致危险错误。当前的安全评估多局限于静态图像、文本或通用风险,难以有效衡量特定场景下的动态危险行为检测能力。为此,我们提出了 HomeSafe-Bench,一个面向家庭场景中危险行为检测的挑战性基准。该基准通过结合物理仿真与先进视频生成技术构建,包含438个多样化案例,覆盖六个功能区域,并提供细粒度的多维标注。除基准外,我们还提出层级双脑家庭安全防护架构(HD-Guard),采用分层流式设计:轻量级 FastBrain 实现高频连续筛查,异步大规模 SlowBrain 执行深层多模态推理,有效平衡推理效率与检测精度。评估结果表明,HD-Guard 在延迟与性能之间取得更优权衡;同时分析揭示了当前基于 VLM 的安全检测存在关键瓶颈。

原文摘要 · Abstract (English)

The rapid evolution of embodied agents has accelerated the deployment of household robots in real-world environments. However, unlike structured industrial settings, household spaces introduce unpredictable safety risks, where system limitations such as perception latency and lack of common sense knowledge can lead to dangerous errors. Current safety evaluations, often restricted to static images, text, or general hazards, fail to adequately benchmark dynamic unsafe action detection in these specific contexts. To bridge this gap, we introduce HomeSafe-Bench, a challenging benchmark designed to evaluate Vision-Language Models (VLMs) on unsafe action detection in household scenarios. HomeSafe-Bench is contrusted via a hybrid pipeline combining physical simulation with advanced video generation and features 438 diverse cases across six functional areas with fine-grained multidimensional annotations. Beyond benchmarking, we propose Hierarchical Dual-Brain Guard for Household Safety (HD-Guard), a hierarchical streaming architecture for real-time safety monitoring. HD-Guard coordinates a lightweight FastBrain for continuous high-frequency screening with an asynchronous large-scale SlowBrain for deep multimodal reasoning, effectively balancing inference efficiency with detection accuracy. Evaluations demonstrate that HD-Guard achieves a superior trade-off between latency and performance, while our analysis identifies critical bottlenecks in current VLM-based safety detection.

视觉语言模型安全检测家庭机器人动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。