arXiv:2510.09771cs.CLcs.AI2025-10中稿 · BLP at AACL-IJCNLP…

用关键词统计与投票机制,低成本实现孟加拉语仇恨言论分类

PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection

  • 通过卡方检验提取关键特征词,结合自适应投票决策
  • 微F1达67.61,优于n-gram基线(60.75)和随机方法(14.65)
  • 适合低资源语言的少样本场景,尤其擅长模糊案例处理

BLP-2025 Task 1A 要求对孟加拉语仇恨言论进行六类分类。传统监督方法依赖大量标注数据,对低资源语言成本高昂。我们提出 PromptGuard,一种少样本框架,结合卡方统计分析提取关键词,并采用自适应多数投票机制做决策。对比了统计关键词选择与随机方法,以及基于共识质量扩展分类轮次的投票机制。卡方关键词在各类别中均带来稳定提升,自适应投票在模糊案例中表现更优。PromptGuard 实现微F1为67.61,高于n-gram基线(60.75)和随机方法(14.65)。消融实验表明,基于卡方的关键词对所有类别均有最一致的正面影响。

原文摘要 · Abstract (English)

The BLP-2025 Task 1A requires Bengali hate speech classification into six categories. Traditional supervised approaches need extensive labeled datasets that are expensive for low-resource languages. We developed PromptGuard, a few-shot framework combining chi-square statistical analysis for keyword extraction with adaptive majority voting for decision-making. We explore statistical keyword selection versus random approaches and adaptive voting mechanisms that extend classification based on consensus quality. Chi-square keywords provide consistent improvements across categories, while adaptive voting benefits ambiguous cases requiring extended classification rounds. PromptGuard achieves a micro-F1 of 67.61, outperforming n-gram baselines (60.75) and random approaches (14.65). Ablation studies confirm chi-square-based keywords show the most consistent impact across all categories.

少样本学习仇恨言论检测孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。