针对数字取证中的仇恨与威胁检测,提出按证据类型选择分析方式的多模态方法。
Detection of Hate and Threat in Digital Forensics: A Case-Driven Multimodal Approach
- 根据文本来源区分嵌入、上下文和纯图像证据,动态选择分析策略。
- 在多源证据场景下表现稳定,推理过程可解释性高。
- 适合需要追查证据链的司法或安全机构使用。
数字取证日益依赖图像、扫描文档和上下文报告等异构证据。这些材料可能包含显性或隐性的伤害、仇恨、威胁、暴力或恐吓内容,但现有自动化方法常假设输入为干净文本,或在无取证依据的情况下应用视觉模型。本文提出一种基于案件驱动的多模态仇恨与威胁检测框架。该框架明确判断文本证据的存在及其来源,区分嵌入式文本、关联上下文文本和仅图像证据。根据识别出的证据配置,框架选择性地应用文本分析、多模态融合或仅图像的语义推理,使用以视觉变换器(ViT)为骨干的视觉语言模型。通过基于证据可用性条件化推理,该方法模拟取证决策流程,提升证据溯源能力,并避免不合理的模态假设。在仿取证图像证据上的实验表明,该方法在异构证据场景下表现出一致且可解释的行为。
原文摘要 · Abstract (English)
Digital forensic investigations increasingly rely on heterogeneous evidence such as images, scanned documents, and contextual reports. These artifacts may contain explicit or implicit expressions of harm, hate, threat, violence, or intimidation, yet existing automated approaches often assume clean text input or apply vision models without forensic justification. This paper presents a case-driven multimodal approach for hate and threat detection in forensic analysis. The proposed framework explicitly determines the presence and source of textual evidence, distinguishing between embedded text, associated contextual text, and image-only evidence. Based on the identified evidence configuration, the framework selectively applies text analysis, multimodal fusion, or image-only semantic reasoning using vision language models with vision transformer backbones (ViT). By conditioning inference on evidence availability, the approach mirrors forensic decision-making, improves evidentiary traceability, and avoids unjustified modality assumptions. Experimental evaluation on forensic-style image evidence demonstrates consistent and interpretable behavior across heterogeneous evidence scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。