融合骨骼与语义模型,实现低延迟隐私保护的公共安全行为检测
From Skeletons to Semantics: Design and Deployment of a Hybrid Edge-Based Action Detection System for Public Safety
- 采用骨骼动作分析与视觉语言模型结合的混合架构
- 在边缘设备上实现低延迟(<100ms)与低资源占用
- 适合对隐私和实时性要求高的公共安全场景
交通枢纽、城市中心和活动场所等公共场所需要及时可靠地检测潜在暴力行为以保障公共安全。尽管自动化视频分析取得显著进展,但在边缘计算条件下,部署仍受限于延迟、隐私和资源消耗。本文提出并部署了一种混合式边缘行为检测系统,结合基于骨骼的动作分析与视觉语言模型进行语义场景理解。骨骼分析支持持续、隐私友好的监控,计算开销低;视觉语言模型则提供上下文理解与零样本推理能力,应对复杂或未见过的情况。本工作不提出新识别模型,而是系统层面比较两种范式在真实边缘约束下的表现。系统在配备GPU的边缘设备上实现,通过演示系统评估延迟、资源使用及运行权衡。结果表明,运动主导与语义主导方法各有优劣,支持选择性地将快速骨骼检测与高层语义推理结合。该系统为公共安全应用中的隐私友好型实时视频分析提供了实用基础。
原文摘要 · Abstract (English)
Public spaces such as transport hubs, city centres, and event venues require timely and reliable detection of potentially violent behaviour to support public safety. While automated video analysis has made significant progress, practical deployment remains constrained by latency, privacy, and resource limitations, particularly under edge-computing conditions. This paper presents the design and demonstrator-based deployment of a hybrid edge-based action detection system that combines skeleton-based motion analysis with vision-language models for semantic scene interpretation. Skeleton-based processing enables continuous, privacy-aware monitoring with low computational overhead, while vision-language models provide contextual understanding and zero-shot reasoning capabilities for complex and previously unseen situations. Rather than proposing new recognition models, the contribution focuses on a system-level comparison of both paradigms under realistic edge constraints. The system is implemented on a GPU-enabled edge device and evaluated with respect to latency, resource usage, and operational trade-offs using a demonstrator-based setup. The results highlight the complementary strengths and limitations of motioncentric and semantic approaches and motivate a hybrid architecture that selectively augments fast skeletonbased detection with higher-level semantic reasoning. The presented system provides a practical foundation for privacy-aware, real-time video analysis in public safety applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。