用知识图谱和大模型蒸馏提升有害梗图识别准确率
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
- 从ConceptNet提取子知识图谱,注入小型视觉语言模型
- 在两个基准数据集上召回率提升35%,优于当前最优方法
- 适合需要高精度识别网络毒害内容的平台安全团队
在线多模态环境中毒性内容识别仍具挑战性,因文本与图像间的上下文关联复杂。本文提出一种新框架,融合大型视觉语言模型(LVLM)的知识蒸馏与知识图谱(KG)注入,增强对仇恨梗图的毒性检测能力。通过从ConceptNet中提取子知识图谱,并嵌入紧凑型视觉语言模型,提升模型对标题中攻击性短语与图像概念之间关系的推理能力。在两个仇恨言论基准数据集上的实验表明,该方法在AU-ROC、F1和召回率上分别优于当前最优基线1.1%、7%和35%。结果验证了结合显式(如知识图谱)与隐式(如LVLM)上下文线索的混合神经符号方法在真实场景中的重要性,有助于实现更准确、可扩展的有毒内容识别,保障网络环境安全。
原文摘要 · Abstract (English)
Toxicity identification in online multimodal environments remains a challenging task due to the complexity of contextual connections across modalities (e.g., textual and visual). In this paper, we propose a novel framework that integrates Knowledge Distillation (KD) from Large Visual Language Models (LVLMs) and knowledge infusion to enhance the performance of toxicity detection in hateful memes. Our approach extracts sub-knowledge graphs from ConceptNet, a large-scale commonsense Knowledge Graph (KG) to be infused within a compact VLM framework. The relational context between toxic phrases in captions and memes, as well as visual concepts in memes enhance the model's reasoning capabilities. Experimental results from our study on two hate speech benchmark datasets demonstrate superior performance over the state-of-the-art baselines across AU-ROC, F1, and Recall with improvements of 1.1%, 7%, and 35%, respectively. Given the contextual complexity of the toxicity detection task, our approach showcases the significance of learning from both explicit (i.e. KG) as well as implicit (i.e. LVLMs) contextual cues incorporated through a hybrid neurosymbolic approach. This is crucial for real-world applications where accurate and scalable recognition of toxic content is critical for creating safer online environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。