构建多模态西语反性别歧视数据集,助力视频内容自动识别性别偏见。
MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videos
- 构建11小时西语社交视频多模态数据集,覆盖文本、语音与视觉信息。
- 发现视觉信息对判断性别歧视至关重要,模型对隐性偏见识别能力弱。
- 适合关注社会偏见检测、多模态分析的研究者使用。
性别歧视指基于性别或性别的偏见与歧视,影响社会制度、人际关系及个体行为。社交媒体通过文本、语音和视觉等多模态传播歧视内容,凸显了多模态分析的必要性。随着短视频平台兴起,性别歧视在视频中广泛传播。自动识别视频中的性别歧视需综合分析语言、音频与视觉元素,极具挑战。本研究提出MuSeD,一个包含约11小时来自TikTok和BitChute的西班牙语视频的多模态数据集;设计创新标注框架,量化文本、语音与视觉在性别歧视分类中的贡献;评估多种大语言模型(LLMs)与多模态大模型在该任务上的表现。结果表明,视觉信息对人类与模型判断性别歧视均起关键作用。模型能有效识别明显性别歧视,但在隐性案例(如刻板印象)上表现不佳,且人工标注者一致性低。这反映了任务内在难度:隐性性别歧视依赖社会文化背景理解。
原文摘要 · Abstract (English)
Sexism is generally defined as prejudice and discrimination based on sex or gender, affecting every sector of society, from social institutions to relationships and individual behavior. Social media platforms amplify the impact of sexism by conveying discriminatory content not only through text but also across multiple modalities, highlighting the critical need for a multimodal approach to the analysis of sexism online. With the rise of social media platforms where users share short videos, sexism is increasingly spreading through video content. Automatically detecting sexism in videos is a challenging task, as it requires analyzing the combination of verbal, audio, and visual elements to identify sexist content. In this study, (1) we introduce MuSeD, a new Multimodal Spanish dataset for Sexism Detection consisting of $\approx$ 11 hours of videos extracted from TikTok and BitChute; (2) we propose an innovative annotation framework for analyzing the contributions of textual, vocal, and visual modalities to the classification of content as either sexist or non-sexist; and (3) we evaluate a range of large language models (LLMs) and multimodal LLMs on the task of sexism detection. We find that visual information plays a key role in labeling sexist content for both humans and models. Models effectively detect explicit sexism; however, they struggle with implicit cases, such as stereotypes, instances where annotators also show low agreement. This highlights the inherent difficulty of the task, as identifying implicit sexism depends on the social and cultural context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。