将执法记录仪视频生成全景摘要,快速理解复杂现场。
Stitching the Story: Creating Panoramic Incident Summaries from Body-Worn Footage
- 用单目SLAM追踪视角,重建环境空间布局。
- 聚类关键视角,选取代表性画面拼成全景图。
- 适合应急响应、事后复盘,提升决策效率。
一线救援人员广泛使用执法记录仪记录事件现场,以支持事后分析。然而,在时间紧迫的情况下,审阅冗长的视频片段不切实际。高效的态势感知需要可快速解读的简洁视觉摘要。本文提出一种计算机视觉流程,将执法记录仪视频转化为能概括事件场景的信息型全景图像。方法基于单目同时定位与地图构建(SLAM)估算相机轨迹,并重建环境空间结构。通过聚类轨迹上的相机位姿识别关键视角,从每个簇中选取代表性帧。利用多帧图像拼接技术将这些帧融合为空间一致的全景图像。生成的摘要可帮助快速理解复杂环境,促进高效决策与事件回溯。
原文摘要 · Abstract (English)
First responders widely adopt body-worn cameras to document incident scenes and support post-event analysis. However, reviewing lengthy video footage is impractical in time-critical situations. Effective situational awareness demands a concise visual summary that can be quickly interpreted. This work presents a computer vision pipeline that transforms body-camera footage into informative panoramic images summarizing the incident scene. Our method leverages monocular Simultaneous Localization and Mapping (SLAM) to estimate camera trajectories and reconstruct the spatial layout of the environment. Key viewpoints are identified by clustering camera poses along the trajectory, and representative frames from each cluster are selected. These frames are fused into spatially coherent panoramic images using multi-frame stitching techniques. The resulting summaries enable rapid understanding of complex environments and facilitate efficient decision-making and incident review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。