用少量文本示例提升图文仇恨言论识别准确率
Bridging Modalities: Enhancing Cross-Modality Hate Speech Detection with Few-Shot In-Context Learning
- 利用大模型少样本上下文学习,实现跨模态知识迁移
- 文本示例使图文仇恨内容识别准确率显著提升
- 文本示范效果优于图文示范,适合快速部署检测系统
互联网上广泛存在的仇恨言论,包括基于文本的推文和图文表情包等形式,对数字平台安全构成严峻挑战。尽管已有研究针对特定模态开发了检测模型,但跨格式检测能力的迁移仍存在明显空白。本研究通过大规模实验,采用大语言模型的少样本上下文学习方法,探索仇恨言论检测在不同模态间的可迁移性。结果表明,基于文本的仇恨言论示例能显著提升图文仇恨内容的分类准确率;且在少样本学习场景中,文本示例的表现优于图文示例。这些发现凸显了跨模态知识迁移的有效性,为优化仇恨言论检测系统提供了重要启示。
原文摘要 · Abstract (English)
The widespread presence of hate speech on the internet, including formats such as text-based tweets and vision-language memes, poses a significant challenge to digital platform safety. Recent research has developed detection models tailored to specific modalities; however, there is a notable gap in transferring detection capabilities across different formats. This study conducts extensive experiments using few-shot in-context learning with large language models to explore the transferability of hate speech detection between modalities. Our findings demonstrate that text-based hate speech examples can significantly enhance the classification accuracy of vision-language hate speech. Moreover, text-based demonstrations outperform vision-language demonstrations in few-shot learning settings. These results highlight the effectiveness of cross-modality knowledge transfer and offer valuable insights for improving hate speech detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。