基于人物交互的可解释暴力威胁检测方法,提升安全场景判断准确性。
Hoi2Threat: An Interpretable Threat Detection Method for Human Violence Scenarios Guided by Human-Object Interaction
- 通过人物交互对(HOI)结构化标签引导语言生成,增强语义理解。
- 在信息正确性、行为映射准确率等指标上比Gemma3(4B)提升7.10%以上。
- 适合需要高可解释性的公共安全监控系统使用。
针对公共安全需求日益增长的背景,自动化威胁检测在高风险场景中的重要性愈发凸显。现有方法普遍存在推理不可解释、语义理解偏差等问题,严重影响实际部署可靠性。为此,本文提出基于人物交互对(HOI-pairs)的威胁检测方法Hoi2Threat,依托细粒度多模态TD-Hoi数据集,利用结构化HOI标签引导语言生成,提升模型对关键实体及其行为交互的语义建模能力。同时设计一套文本响应质量评估指标,系统衡量模型在威胁解释过程中的表征准确性与可读性。实验表明,Hoi2Threat在多个威胁检测任务中显著提升,尤其在信息正确性(CoI)、行为映射准确率(BMA)和威胁细节定位(TDO)三项核心指标上,相较Gemma3(4B)分别提升7.10%、6.80%和2.63%。
原文摘要 · Abstract (English)
In light of the mounting imperative for public security, the necessity for automated threat detection in high-risk scenarios is becoming increasingly pressing. However, existing methods generally suffer from the problems of uninterpretable inference and biased semantic understanding, which severely limits their reliability in practical deployment. In order to address the aforementioned challenges, this article proposes a threat detection method based on human-object interaction pairs (HOI-pairs), Hoi2Threat. This method is based on the fine-grained multimodal TD-Hoi dataset, enhancing the model's semantic modeling ability for key entities and their behavioral interactions by using structured HOI tags to guide language generation. Furthermore, a set of metrics is designed for the evaluation of text response quality, with the objective of systematically measuring the model's representation accuracy and comprehensibility during threat interpretation. The experimental results have demonstrated that Hoi2Threat attains substantial enhancement in several threat detection tasks, particularly in the core metrics of Correctness of Information (CoI), Behavioral Mapping Accuracy (BMA), and Threat Detailed Orientation (TDO), which are 5.08, 5.04, and 4.76, and 7.10%, 6.80%, and 2.63%, respectively, in comparison with the Gemma3 (4B). The aforementioned results provide comprehensive validation of the merits of this approach in the domains of semantic understanding, entity behavior mapping, and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。