用开源嵌入模型精准识别德语新闻评论中的性别歧视。
Detecting Sexism in German Online Newspaper Comments with Open-Source Text Embeddings (Team GDA, GermEval2024 Shared Task 1: GerMS-Detect, Subtasks 1 and 2, Closed Track)
- 基于开源文本嵌入训练分类器,模拟人类标注者判断。
- 在德语反性别歧视任务中获4th名,平均宏F1达0.597。
- 计算高效,适合多语言场景的规模化部署。
在线媒体评论中的性别歧视问题普遍存在且常以微妙形式出现,导致内容审核困难,因对何为性别歧视的理解因人而异。本文研究单语和多语开源文本嵌入,用于可靠检测奥地利某报纸德语评论中的性别歧视与厌女言论。基于文本嵌入训练的分类器能紧密模仿人类标注者的个体判断。该方法在GermEval 2024 GerMS-Detect Subtask 1挑战中表现稳健,平均宏F1得分为0.597(在Codabench上排名第4)。同时,在Subtask 2中准确预测了人类标注分布,平均杰恩-申诺尔距离为0.301(排名第2)。其计算效率表明,该方法具备在多种语言和语言情境下进行可扩展应用的潜力。
原文摘要 · Abstract (English)
Sexism in online media comments is a pervasive challenge that often manifests subtly, complicating moderation efforts as interpretations of what constitutes sexism can vary among individuals. We study monolingual and multilingual open-source text embeddings to reliably detect sexism and misogyny in German-language online comments from an Austrian newspaper. We observed classifiers trained on text embeddings to mimic closely the individual judgements of human annotators. Our method showed robust performance in the GermEval 2024 GerMS-Detect Subtask 1 challenge, achieving an average macro F1 score of 0.597 (4th place, as reported on Codabench). It also accurately predicted the distribution of human annotations in GerMS-Detect Subtask 2, with an average Jensen-Shannon distance of 0.301 (2nd place). The computational efficiency of our approach suggests potential for scalable applications across various languages and linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。