构建多模态细粒度反女性主义数据集,提升社交媒体视频中隐性性别歧视识别能力。
Beyond Binary Classification: Detecting Fine-Grained Sexism in Social Media Videos
- 构建西班牙语多模态数据集FineMuSe,支持二分类与细粒度标注
- 提出包含讽刺、幽默等修辞手法的分层分类体系
- 大模型在识别复杂性别歧视时表现接近人工,但视觉线索联合判断仍有不足
网络性别歧视形式多样,传统自动化工具多局限于二分类,难以捕捉细微表现。为此,本文提出FineMuSe数据集,覆盖西班牙语社交媒体视频,包含二分类与细粒度标注;构建涵盖性别歧视类型、非性别歧视内容及讽刺、幽默等修辞手法的分层分类体系;评估多种大语言模型在二分类与细粒度检测任务上的表现。结果表明,多模态大模型在识别复杂性别歧视方面表现媲美人类标注者,但在通过视觉线索传递的多重歧视共现场景下仍存在识别困难。
原文摘要 · Abstract (English)
Online sexism appears in various forms, which makes its detection challenging. Although automated tools can enhance the identification of sexist content, they are often restricted to binary classification. Consequently, more subtle manifestations of sexism may remain undetected due to the lack of fine-grained, context-sensitive labels. To address this issue, we make the following contributions: (1) we present FineMuSe, a new multimodal sexism detection dataset in Spanish that includes both binary and fine-grained annotations; (2) we introduce a comprehensive hierarchical taxonomy that encompasses forms of sexism, non-sexism, and rhetorical devices of irony and humor; and (3) we evaluate a wide range of LLMs for both binary and fine-grained sexism detection. Our findings indicate that multimodal LLMs perform competitively with human annotators in identifying nuanced forms of sexism; however, they struggle to capture co-occurring sexist types when these are conveyed through visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。