arXiv:2503.21504cs.CLcs.AI2025-03被引 1

用多模态数据识别隐晦用语,提升网络内容审核准确率

Keyword-Oriented Multimodal Modeling for Euphemism Identification

  • 以关键词为核心构建图文音多模态识别框架
  • 在毒品、武器、性相关数据集上显著超越现有模型
  • 适合内容安全、隐晦表达检测方向的研究者

隐晦用语识别旨在揭示隐喻性表述的真实含义,例如将非法文本中的“weed”(隐晦用语)关联到“marijuana”(目标关键词),有助于内容审核和打击地下市场。现有方法多基于文本,但社交媒体兴起促使需融合文本、图像和音频的多模态分析。然而,缺乏针对隐晦用语的多模态数据集制约了研究进展。为此,本文首次将隐晦用语及其对应的目标关键词视为关键词,构建了关键词导向的多模态隐晦用语语料库(KOM-Euph),包含毒品(Drug)、武器(Weapon)和性相关(Sexuality)三类数据集,涵盖文本、图像和语音。同时提出关键词导向的多模态隐晦用语识别方法(KOM-EI),通过跨模态特征对齐与动态融合模块,显式利用关键词的视觉和音频特征,实现高效识别。大量实验表明,KOM-EI优于当前最优模型及大语言模型,并验证了所建多模态数据集的有效性。

原文摘要 · Abstract (English)

Euphemism identification deciphers the true meaning of euphemisms, such as linking "weed" (euphemism) to "marijuana" (target keyword) in illicit texts, aiding content moderation and combating underground markets. While existing methods are primarily text-based, the rise of social media highlights the need for multimodal analysis, incorporating text, images, and audio. However, the lack of multimodal datasets for euphemisms limits further research. To address this, we regard euphemisms and their corresponding target keywords as keywords and first introduce a keyword-oriented multimodal corpus of euphemisms (KOM-Euph), involving three datasets (Drug, Weapon, and Sexuality), including text, images, and speech. We further propose a keyword-oriented multimodal euphemism identification method (KOM-EI), which uses cross-modal feature alignment and dynamic fusion modules to explicitly utilize the visual and audio features of the keywords for efficient euphemism identification. Extensive experiments demonstrate that KOM-EI outperforms state-of-the-art models and large language models, and show the importance of our multimodal datasets.

隐晦用语多模态内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。