小模型通过优化提示和数据增强,也能高效识别隐晦的仇恨内容。
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- 用结构化提示和细粒度标签提升小模型性能
- 自动生成中性反事实图像,减少误判率
- 适合资源有限但需高鲁棒性的部署场景
在线仇恨内容仍是重大社会挑战,尤其在多模态内容中,仇恨常以文本与图像交互、幽默等隐晦方式呈现,难以被自动化系统识别。尽管当前视觉-语言模型(VLMs)对显性仇恨表现良好,但其部署受限于高推理成本及对细微内容的识别失败。本文研究如何通过提示优化、微调和自动数据增强,提升小型模型的表现。提出一个端到端流程,调整提示结构、标签粒度和训练模态,发现结构化提示与规模化监督可显著增强紧凑型VLM。同时构建多模态增强框架,通过协调使用的LLM-VLM生成反事实中性漫画,降低虚假相关性,提升对隐性仇恨的检测能力。消融实验量化各组件贡献,表明提示设计、细粒度标签与针对性增强共同缩小了小模型与大模型之间的差距。结果为无需依赖昂贵大模型推理的更鲁棒、可部署的多模态仇恨检测系统提供了可行路径。
原文摘要 · Abstract (English)
Online hate remains a significant societal challenge, especially as multimodal content enables subtle, culturally grounded, and implicit forms of harm. Hateful memes embed hostility through text-image interactions and humor, making them difficult for automated systems to interpret. Although recent Vision-Language Models (VLMs) perform well on explicit cases, their deployment is limited by high inference costs and persistent failures on nuanced content. This work examines how far small models can be improved through prompt optimization, fine-tuning, and automated data augmentation. We introduce an end-to-end pipeline that varies prompt structure, label granularity, and training modality, showing that structured prompts and scaled supervision significantly strengthen compact VLMs. We also develop a multimodal augmentation framework that generates counterfactually neutral memes via a coordinated LLM-VLM setup, reducing spurious correlations and improving the detection of implicit hate. Ablation studies quantify the contribution of each component, demonstrating that prompt design, granular labels, and targeted augmentation collectively narrow the gap between small and large models. The results offer a practical path toward more robust and deployable multimodal hate-detection systems without relying on costly large-model inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。