arXiv:2411.10480cs.CVcs.AI2024-11被引 1

通过上下文提示与细粒度标注提升仇恨梗图识别准确率

Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling

  • 设计端到端优化框架,融合模态、提示词与标注策略
  • 在Hateful Meme数据集上达到最高准确率与AUROC
  • 验证了提示词与标注细节对模型性能的关键作用

社交媒体上多模态内容泛滥,给自动化内容审核带来挑战。这需要提升多模态分类能力,并深入理解图像与梗图中的隐含含义。尽管已有研究通过微调提升模型性能,但很少有工作探索涵盖模态、提示词、标注和微调的端到端优化流程。本文提出一种面向复杂任务的端到端概念性优化框架。实验表明该传统但新颖的框架有效,显著提升了准确率与AUROC指标。消融实验显示,单独优化提示词或标注也具有实际效果。

原文摘要 · Abstract (English)

The prevalence of multi-modal content on social media complicates automated moderation strategies. This calls for an enhancement in multi-modal classification and a deeper understanding of understated meanings in images and memes. Although previous efforts have aimed at improving model performance through fine-tuning, few have explored an end-to-end optimization pipeline that accounts for modalities, prompting, labeling, and fine-tuning. In this study, we propose an end-to-end conceptual framework for model optimization in complex tasks. Experiments support the efficacy of this traditional yet novel framework, achieving the highest accuracy and AUROC. Ablation experiments demonstrate that isolated optimizations are not ineffective on their own.

多模态仇恨内容检测提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。