通过多模态因果推理,提升低噪环境下战术视频的威胁识别准确率
A Tactical Behaviour Recognition Framework Based on Causal Multimodal Reasoning: A Study on Covert Audio-Video Analysis Combining GAN Structure Enhancement and Phonetic Accent Modelling
- 融合视觉、音频与动作线索构建时序图,利用图注意力分析跨模态关联
- 在噪声环境中实现89.3%的时间对齐准确率和超85%的完整威胁链识别率
- 适用于安防监控、军事防御等需高精度实时威胁感知的场景
本文提出TACTIC-GRAPHS系统,结合谱图理论与多模态图神经推理,实现高噪声、弱结构下战术视频的语义理解与威胁检测。该框架引入谱嵌入、时序因果边建模及跨异构模态的判别路径推理,采用语义感知关键帧提取方法融合视觉、声学与动作线索构建时序图。通过图注意力与拉普拉斯谱映射,实现跨模态加权与因果信号分析。在TACTIC-AVS与TACTIC-Voice数据集上的实验表明,系统在时间对齐任务中达到89.3%准确率,完整威胁链识别超过85%,节点延迟控制在±150毫秒内。该方法提升了结构可解释性,支持安防、国防与智能安全系统的应用。
原文摘要 · Abstract (English)
This paper introduces TACTIC-GRAPHS, a system that combines spectral graph theory and multimodal graph neural reasoning for semantic understanding and threat detection in tactical video under high noise and weak structure. The framework incorporates spectral embedding, temporal causal edge modeling, and discriminative path inference across heterogeneous modalities. A semantic-aware keyframe extraction method fuses visual, acoustic, and action cues to construct temporal graphs. Using graph attention and Laplacian spectral mapping, the model performs cross-modal weighting and causal signal analysis. Experiments on TACTIC-AVS and TACTIC-Voice datasets show 89.3 percent accuracy in temporal alignment and over 85 percent recognition of complete threat chains, with node latency within plus-minus 150 milliseconds. The approach enhances structural interpretability and supports applications in surveillance, defense, and intelligent security systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。