用软标签优化模型,提升表情包性别歧视意图识别效果
BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes
- 基于XLM-RoBERTa的文本方法,融合软标签与硬标签损失函数
- 软标签指标得分0.3229,硬标签F1达0.4236,排名前半
- 揭示标注分歧对主观任务建模的关键影响,适合情感分析研究者
本文介绍BioSentinel团队在CLEF 2026评估活动的EXIST 2026任务2.2:表情包源意图分类中的参与情况。该任务要求在学习分歧(Le-Wi-Di)范式下,将表情包的交际意图分类为直接、评判性或无(非性别歧视)。模型需同时输出硬标签和软标签(概率分布)。我们提出一种以xlm-roberta-base(270M参数)为基础的文本中心方法,采用复合损失函数,结合软标注分布上的KL散度与硬标签上的加权交叉熵。在官方测试集上,系统获得ICM-Soft-Norm 0.3229、ICM-Norm 0.3778,硬标签F1-score为0.4236,在软-软评估中排名第40位(共118项提交),在硬-硬评估中排名第49位(共187项提交)。我们分析了数据集特性、大模型架构的探索性实验,以及标注分歧对主观自然语言处理任务建模范式的影响。消融实验表明,KL损失提升软标签性能,交叉熵损失增强硬标签准确率。此外,还报告了验证集温度调节分析结果。
原文摘要 · Abstract (English)
This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (Le-Wi-Di) paradigm that mandates both hard-label and soft-label (probability distribution) predictions. We present a text-centric approach built on xlm-roberta-base (270M parameters) trained with a composite loss function combining KL divergence on soft annotator distributions and weighted cross-entropy on hard labels. On the official test set, the system achieved an ICM-Soft-Norm of 0.3229 and ICM-Norm of 0.3778, with a hard F1-score of 0.4236, ranking 40th (out of 118 submissions) in the soft-soft evaluation and 49th (out of 187 submissions) in the hard-hard evaluation. We provide an analysis of the dataset characteristics, exploratory larger-architecture runs, and the role of annotator disagreement in shaping model design for subjective NLP tasks. Ablation results show that KL loss improves soft-label metrics, while CE loss improves hard-label accuracy. We also report a separate validation-set temperature analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。