用多模态大模型识别新加坡语境下的攻击性迷因,提升内容审核精准度。
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
- 基于GPT-4V标注11.2万张迷因,微调70亿参数多模态模型
- 在测试集上达到80.62%准确率和0.8192 AUROC
- 整合OCR、翻译与视觉语言模型,适合本地化内容审核
传统在线内容审核系统难以分类现代多模态表达形式,如迷因——这种高度复杂且信息密集的媒介。在新加坡这样文化多元、低资源语言广泛使用的社会中,理解网络内容需依赖大量本地知识。本文构建了一个包含11.2万张迷因的大规模数据集,由GPT-4V进行标注,用于微调视觉语言模型(VLM)以识别新加坡语境下的攻击性迷因。提出一个包含OCR、翻译与70亿参数分类器的处理流程,该方案在保留测试集上实现80.62%准确率和0.8192 AUROC,显著辅助人工审核。相关数据集、代码与模型权重已开源至https://github.com/aliencaocao/vlm-for-memes-aisg。
原文摘要 · Abstract (English)
Traditional online content moderation systems struggle to classify modern multimodal means of communication, such as memes, a highly nuanced and information-dense medium. This task is especially hard in a culturally diverse society like Singapore, where low-resource languages are used and extensive knowledge on local context is needed to interpret online content. We curate a large collection of 112K memes labeled by GPT-4V for fine-tuning a VLM to classify offensive memes in Singapore context. We show the effectiveness of fine-tuned VLMs on our dataset, and propose a pipeline containing OCR, translation and a 7-billion parameter-class VLM. Our solutions reach 80.62% accuracy and 0.8192 AUROC on a held-out test set, and can greatly aid human in moderating online contents. The dataset, code, and model weights have been open-sourced at https://github.com/aliencaocao/vlm-for-memes-aisg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。