用提问式弱监督分解识别梗图仇恨言论,多语言表现更优。
MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection

- 将梗图理解拆解为问答任务,限制答案选项提升准确性。
- 在英、中、印语种上分别获第1、2、3名,中文和印地语提升显著。
- 适合多语言、跨文化仇恨言论检测场景,尤其对隐含讽刺有效。
由于梗图的多模态特性及隐含的文化语境(如反讽、背景依赖),其仇恨言论检测极具挑战性。尽管视觉-语言模型(VLM)能联合分析图文信息,但端到端提示往往脆弱——单次预测需同时判断目标、立场、隐晦性与讽刺性,易出错,多语言环境下问题更突出。本文提出提问式弱监督(PWS)方法,将梗图理解分解为针对性的问答标签函数,限定答案范围,专攻同性恋与跨性别歧视检测。借助量化版Qwen3-VLM回答特定问题提取特征,相比直接分类,性能显著提升,尤其在中文与印地语中优势明显:英文排名首位,中文第二,印地语第三。通过错误驱动的标签函数迭代扩展与特征剪枝,减少冗余并增强泛化能力。结果表明,提问式弱监督在多语言多模态仇恨言论检测中具有高度有效性。
原文摘要 · Abstract (English)
Detecting hate speech in memes is challenging due to their multimodal nature and subtle, culturally grounded cues such as sarcasm and context. While recent vision-language models (VLMs) enable joint reasoning over text and images, end-to-end prompting can be brittle, as a single prediction must resolve target, stance, implicitness, and irony. These challenges are amplified in multilingual settings. We propose a prompted weak supervision (PWS) approach that decomposes meme understanding into targeted, question-based labeling functions with constrained answer options for homophobia and transphobia detection in the LT-EDI 2026 shared task. Using a quantized Qwen3-VLM to extract features by answering targeted questions, our method outperforms direct VLM classification, with substantial gains for Chinese and Hindi, ranking 1st in English, 2nd in Chinese, and 3rd in Hindi. Iterative refinement via error-driven LF expansion and feature pruning reduces redundancy and improves generalization. Our results highlight the effectiveness of prompted weak supervision for multilingual multimodal hate speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。