用 memes 数据辅助训练仇恨视频检测模型,解决数据少难题。
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
- 用 meme 数据替代或扩充视频数据进行训练。
- 在资源稀缺时性能超越现有基准,提升明显。
- 适合关注视频内容安全与跨模态学习的研究者。
在线内容中的仇恨言论检测对构建安全数字空间至关重要。尽管文本和 meme 模态已取得显著进展,但基于视频的仇恨言论检测仍因标注数据匮乏和视频标注成本高而研究不足。这一问题在大模型依赖海量训练数据的背景下尤为突出。为此,本文提出利用 meme 数据集作为视频数据的替代与增强策略。通过设计人工辅助的重新标注流程,将 meme 数据集标签与视频数据集对齐,仅需少量标注工作即可保证一致性。使用两种先进的视觉-语言模型验证,meme 数据可在资源有限场景下替代视频数据,并进一步扩充视频数据集以提升性能。实验结果持续优于现有最佳基准,证明了跨模态迁移学习在推进仇恨视频检测方面的潜力。数据集与代码已公开于 https://github.com/Social-AI-Studio/CrossModalTransferLearning。
原文摘要 · Abstract (English)
Detecting hate speech in online content is essential to ensuring safer digital spaces. While significant progress has been made in text and meme modalities, video-based hate speech detection remains under-explored, hindered by a lack of annotated datasets and the high cost of video annotation. This gap is particularly problematic given the growing reliance on large models, which demand substantial amounts of training data. To address this challenge, we leverage meme datasets as both a substitution and an augmentation strategy for training hateful video detection models. Our approach introduces a human-assisted reannotation pipeline to align meme dataset labels with video datasets, ensuring consistency with minimal labeling effort. Using two state-of-the-art vision-language models, we demonstrate that meme data can substitute for video data in resource-scarce scenarios and augment video datasets to achieve further performance gains. Our results consistently outperform state-of-the-art benchmarks, showcasing the potential of cross-modal transfer learning for advancing hateful video detection. Dataset and code are available at https://github.com/Social-AI-Studio/CrossModalTransferLearning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。