arXiv:2509.18571cs.CV2025-09

实时视频威胁监测框架,兼顾速度与可解释性。

Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought

  • 将视频帧分解为人物-物体-交互-地点语义元组,压缩信息损失。
  • 在线去重更新机制降低冗余,实现毫秒级响应。
  • 用思维链微调大模型,生成逻辑清晰的威胁评估报告。

实时威胁监测需在视频流中识别威胁行为,并通过解释性文本提供事件推理与评估。然而,现有基于监督学习或生成模型的方法难以同时满足实时性能与决策可解释性的高要求。为此,我们提出Live-E2T框架,通过三项协同机制统一双重目标:首先,将视频帧解构为结构化的人-物-交互-地点语义元组,形成紧凑且语义聚焦的表征,避免传统特征压缩中的信息退化;其次,提出高效的在线事件去重与更新机制,过滤时空冗余,保障系统实时响应;最后,采用思维链(Chain-of-Thought)策略微调大型语言模型,赋予其对事件序列进行透明、逻辑性强的推理能力,生成连贯的威胁评估报告。在XD-Violence和UCF-Crime等基准数据集上的大量实验表明,Live-E2T在威胁检测准确率、实时效率及关键的可解释性维度上均显著优于当前最优方法。

原文摘要 · Abstract (English)

Real-time threat monitoring identifies threatening behaviors in video streams and provides reasoning and assessment of threat events through explanatory text. However, prevailing methodologies, whether based on supervised learning or generative models, struggle to concurrently satisfy the demanding requirements of real-time performance and decision explainability. To bridge this gap, we introduce Live-E2T, a novel framework that unifies these two objectives through three synergistic mechanisms. First, we deconstruct video frames into structured Human-Object-Interaction-Place semantic tuples. This approach creates a compact, semantically focused representation, circumventing the information degradation common in conventional feature compression. Second, an efficient online event deduplication and updating mechanism is proposed to filter spatio-temporal redundancies, ensuring the system's real time responsiveness. Finally, we fine-tune a Large Language Model using a Chain-of-Thought strategy, endow it with the capability for transparent and logical reasoning over event sequences to produce coherent threat assessment reports. Extensive experiments on benchmark datasets, including XD-Violence and UCF-Crime, demonstrate that Live-E2T significantly outperforms state-of-the-art methods in terms of threat detection accuracy, real-time efficiency, and the crucial dimension of explainability.

视频监控威胁检测可解释AI大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。