arXiv:2507.20924cs.CLcs.AI2025-07被引 1

用可解释的词汇瓶颈模型检测社交媒体中的性别歧视,兼顾准确与透明度。

FHSTP@EXIST 2025 Benchmark: Sexism Detection with Transparent Speech Concept Bottleneck Models

  • 用形容词作为可解释概念,通过大模型生成语义表示
  • 在多语言数据上取得6~7名的竞争力结果,支持细粒度解释
  • 适合关注公平性、可解释AI的研究者和内容安全从业者

性别歧视在社交媒体和在线对话中日益普遍。为应对这一问题,CLEF 2025启动了第五届社交网络性别歧视识别(EXIST)挑战赛。本文聚焦于该赛事首个任务:识别并分类社交媒体文本中的性别歧视。我们针对三个子任务提出解决方案并报告结果:1.1——推文中的性别歧视识别,1.2——推文来源意图判断,1.3——推文性别歧视类别分类。我们构建了三种模型分别应对各子任务,形成三次独立运行:语音概念瓶颈模型(SCBM)、融合Transformer的语音概念瓶颈模型(SCBMT),以及微调后的XLM-RoBERTa模型。SCBM使用描述性形容词作为人类可理解的概念瓶颈,利用大语言模型(LLMs)将输入文本编码为形容词表示,再训练轻量级分类器完成下游任务。SCBMT在SCBM基础上融合形容词表示与Transformer上下文嵌入,平衡可解释性与分类性能。此外,我们还探索了利用标注者人口统计信息等元数据的潜力。在1.1任务中,微调后的XLM-RoBERTa在英语和西班牙语上分别获得第6名,在英语软-软评估中排名第4;我们的SCBMT在英语和西班牙语上分别位列第7和第6名。

原文摘要 · Abstract (English)

Sexism has become widespread on social media and in online conversation. To help address this issue, the fifth Sexism Identification in Social Networks (EXIST) challenge is initiated at CLEF 2025. Among this year's international benchmarks, we concentrate on solving the first task aiming to identify and classify sexism in social media textual posts. In this paper, we describe our solutions and report results for three subtasks: Subtask 1.1 - Sexism Identification in Tweets, Subtask 1.2 - Source Intention in Tweets, and Subtask 1.3 - Sexism Categorization in Tweets. We implement three models to address each subtask which constitute three individual runs: Speech Concept Bottleneck Model (SCBM), Speech Concept Bottleneck Model with Transformer (SCBMT), and a fine-tuned XLM-RoBERTa transformer model. SCBM uses descriptive adjectives as human-interpretable bottleneck concepts. SCBM leverages large language models (LLMs) to encode input texts into a human-interpretable representation of adjectives, then used to train a lightweight classifier for downstream tasks. SCBMT extends SCBM by fusing adjective-based representation with contextual embeddings from transformers to balance interpretability and classification performance. Beyond competitive results, these two models offer fine-grained explanations at both instance (local) and class (global) levels. We also investigate how additional metadata, e.g., annotators' demographic profiles, can be leveraged. For Subtask 1.1, XLM-RoBERTa, fine-tuned on provided data augmented with prior datasets, ranks 6th for English and Spanish and 4th for English in the Soft-Soft evaluation. Our SCBMT achieves 7th for English and Spanish and 6th for Spanish.

性别歧视检测可解释AI概念瓶颈多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。