将执法记录仪视频拆解为10秒片段,自动标记情境与动作强度,提升训练效率。
Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC

- 将视频切分为10秒窗口,用CLIP编码帧并聚合表示
- 情境分类准确率78.75%,动作强度识别达88.33%
- 支持快速审案与训练,隐私保护设计可复用
执法机构积累大量执法记录仪(BWC)视频,但其操作过程仍不透明。分析人员需长时间观看完整视频以定位关键事件起始点及活动强度变化点。本文提出一种处理BWC视频的方法:将视频转换为时间对齐的固定长度10秒窗口,并采用隐私保护协议进行标注。每个窗口标注两个维度信息:(i)操作情境;(ii)窗口内动作强度,对因黑暗、模糊或遮挡导致证据不足的窗口标注为低置信度标签。使用从每窗口采样帧经CLIP模型编码后聚合形成窗口级表示,结合密集光流统计量捕捉运动强度。在测试集上,最佳情境分类模型准确率达78.75%,最佳动作模型准确率为88.33%。通过完整性审计验证结果,证明视觉时间线有助于加速事件审查,并使警员训练流程更高效实用。
原文摘要 · Abstract (English)
Law enforcement agencies are accumulating vast amounts of body-worn camera (BWC) footage. However, this remains operationally opaque. That is, analysts and trainers still have to invest considerable time watching full-length videos to pinpoint the start of key encounters and identify the points where activity shifts to something more physically intense. We present an approach to process BWC video into a time-aligned sequence of fixed-length 10-second windows, processed and labeled using a privacy-conscious protocol. Each window is labeled with two dimensions of information: (i) the operational context of the window and (ii) the level of motion intensity within the window, with low-evidence labels for windows for which insufficient evidence exists due to darkness, blur or occlusion. We train models to classify windows based on these two axes using frames sampled from each window encoded using CLIP model and aggregated into a window-level representation. We extract dense optical flow statistics for each window to capture motion intensity. On test windows the best context model achieves 78.75% accuracy, and the best-accuracy activity model achieves 88.33%. We also included integrity audits to show the results and how the visual timeline representations support faster incident review and make the officer training workflow more practical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。