用大模型检测隐性歧视语言,识别对弱势群体的居高临下式话语。
PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
- 基于双语数据集训练大模型,捕捉隐性贬损语气。
- 发现不同弱势群体受偏见影响程度差异显著。
- 适合关注网络伦理与公平性的研究者使用。
免责声明:本文样本可能具有危害性并引发不适!居高临下与轻蔑言语(PCL)是一种针对弱势群体的言语形式,作为毒性语言的重要分支,此类语言加剧了网络社区间的冲突,对弱势群体造成负面影响。传统预训练语言模型因难以识别其隐性毒性特征(如虚伪同情、表里不一)而表现不佳。随着大语言模型(LLMs)的发展,可利用其丰富的语义情感信息构建隐性毒性检测新范式。本文提出PclGPT,一个专为PCL设计的综合性大模型基准。通过收集、标注并整合Pcl-PT/SFT数据集,采用系统的预训练与监督微调阶梯流程,开发出双语版本PclGPT-EN/CN模型组,以支持隐性毒性检测。群体检测结果及细粒度分析显示,不同弱势群体在受到PCL偏见影响的程度上存在显著差异,亟需社会更多关注以保护其权益。
原文摘要 · Abstract (English)
Disclaimer: Samples in this paper may be harmful and cause discomfort! Patronizing and condescending language (PCL) is a form of speech directed at vulnerable groups. As an essential branch of toxic language, this type of language exacerbates conflicts and confrontations among Internet communities and detrimentally impacts disadvantaged groups. Traditional pre-trained language models (PLMs) perform poorly in detecting PCL due to its implicit toxicity traits like hypocrisy and false sympathy. With the rise of large language models (LLMs), we can harness their rich emotional semantics to establish a paradigm for exploring implicit toxicity. In this paper, we introduce PclGPT, a comprehensive LLM benchmark designed specifically for PCL. We collect, annotate, and integrate the Pcl-PT/SFT dataset, and then develop a bilingual PclGPT-EN/CN model group through a comprehensive pre-training and supervised fine-tuning staircase process to facilitate implicit toxic detection. Group detection results and fine-grained detection from PclGPT and other models reveal significant variations in the degree of bias in PCL towards different vulnerable groups, necessitating increased societal attention to protect them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。