首个中文多模态贬低性语言数据集,助力识别隐性歧视言语与表情。
Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector
- 构建715个B站视频的多模态标注数据集,包含高精度面部表情片段。
- 提出MultiPCL检测器,融合表情分析显著提升贬低性语言识别效果。
- 适合研究网络歧视、多模态内容安全及中文社会语义理解的学者。
贬低与居高临下的语言(PCL)是一种针对弱势群体的歧视性有毒言论,威胁线上线下安全。尽管毒性言论研究多聚焦于明显的仇恨言论,但以微侵略形式存在的PCL仍缺乏关注。此外,主流群体对弱势社群表现出的歧视性面部表情和态度,其影响可能超过语言本身,但此类视觉特征常被忽视。本文首次引入PCLMM数据集,这是首个针对中文场景的多模态PCL数据集,包含来自Bilibili的715个已标注视频,涵盖高质量的PCL面部帧片段。同时提出MultiPCL检测器,集成面部表情检测模块,验证了多模态互补在该挑战性任务中的有效性。本工作为有毒言论领域中微侵略检测的进展提供了重要贡献。
原文摘要 · Abstract (English)
Patronizing and Condescending Language (PCL) is a form of discriminatory toxic speech targeting vulnerable groups, threatening both online and offline safety. While toxic speech research has mainly focused on overt toxicity, such as hate speech, microaggressions in the form of PCL remain underexplored. Additionally, dominant groups' discriminatory facial expressions and attitudes toward vulnerable communities can be more impactful than verbal cues, yet these frame features are often overlooked. In this paper, we introduce the PCLMM dataset, the first Chinese multimodal dataset for PCL, consisting of 715 annotated videos from Bilibili, with high-quality PCL facial frame spans. We also propose the MultiPCL detector, featuring a facial expression detection module for PCL recognition, demonstrating the effectiveness of modality complementarity in this challenging task. Our work makes an important contribution to advancing microaggression detection within the domain of toxic speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。