arXiv:2605.24344cs.CL2026-05

破解中文有害梗图识别难题,通过文化知识增强实现可解释判断

Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes

论文配图:Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
图 1 · 摘自论文原文
  • 构建首个中文有害梗图解释数据集,提供正反双重视角
  • 引入文化知识库与相对意图推理框架,提升模糊内容判别力
  • 开源数据集与代码,助力中文网络内容安全研究

有害梗图检测研究备受关注,已涌现出众多数据集与方法。然而,中文有害梗图检测进展滞后,主要受两大挑战制约:其一,准确评估梗图危害性高度依赖深层文化背景理解;其二,许多梗图语义模糊,危害性判定极具主观性。为此,本文聚焦可解释的中文有害梗图检测,构建首个中文有害梗图解释数据集 Ex-ToxiCN-MM。该数据集为每张梗图提供“有害”与“无害”两类对立解释,旨在严格评估模型对文化语境下模糊内容的辨别与理解能力。我们还构建了专门的中文文化概念与辱骂词汇知识库(C-HarmKB),为模型提供关键先验知识。针对梗图歧义性与背景知识缺失问题,提出综合归因分析框架 RIKE,包含归因知识增强模块(AKE)与相对意图推理模块(RIR)。大量定量与定性实验表明,所提方法在多项指标上超越主流基线模型,在中文有害梗图归因任务中表现更优。本研究相关代码、Ex-ToxiCN-MM 数据集及 C-HarmKB 知识库已开源至 https://github.com/wimiw123/Ex-ToxiCN-MM。

原文摘要 · Abstract (English)

Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful memes lags considerably, primarily due to two challenges: first, accurately assessing a meme's harmfulness depends heavily on understanding deep cultural context; second, many memes are semantically ambiguous, making harmfulness highly subjective. To address these issues, we focus on the interpretable detection of Chinese harmful memes by constructing the first Chinese harmful meme explanation dataset, Ex-ToxiCN-MM. This dataset offers opposing interpretations, categorized as "harmful" and "non-harmful", for each meme, aiming to rigorously evaluate a model's ability to discern and comprehend ambiguous, culturally grounded content. We built a specialized knowledge base of Chinese cultural concepts and offensive vocabulary to supply models with essential prior knowledge (C-HarmKB). To address the ambiguity and lack of background knowledge in meme attribution, we have developed a comprehensive attribution analysis framework, RIKE, which includes an Attribution Knowledge Enhancement module (AKE) and a Relative Intent Reasoning module (RIR). Extensive quantitative and qualitative experiments demonstrate that our method outperforms mainstream baseline models across multiple metrics in the task of attributing harmful memes in Chinese. The code, Ex-ToxiCN-MM dataset, and Chinese Harmful Semantic Knowledge Base (C-HarmKB) involved in this study have been open-sourced at https://github.com/wimiw123/Ex-ToxiCN-MM

有害内容检测文化语境可解释性知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。