用人文研究指导标注,提升毒性语言识别的公平性与多样性。
Mitigating Biases to Embrace Diversity: A Comprehensive Annotation Benchmark for Toxic Language
- 基于人文研究设计结构化标注指南,减少主观偏差。
- 新数据集使人类与大模型标注一致性显著提高。
- 小模型用多源大模型标注数据训练,效果优于大人工数据集。
本研究提出一种基于人文研究的规范性标注基准,以确保对攻击性语言(尤其非主流和日常用语)标注的一致性与无偏性。我们构建了两个新标注数据集,其人类与大语言模型(LLM)标注的一致性高于基于描述性指令的原始数据集。实验表明,在缺乏专业标注员时,大语言模型可作为有效替代方案。此外,基于多源大模型标注数据微调的小模型,性能优于在更大、单一来源人工标注数据上训练的模型。这些发现凸显了结构化指南在降低主观差异、有限数据下保持性能以及包容语言多样性方面的价值。内容警告:本文仅用于学术分析攻击性语言,需谨慎阅读。
原文摘要 · Abstract (English)
This study introduces a prescriptive annotation benchmark grounded in humanities research to ensure consistent, unbiased labeling of offensive language, particularly for casual and non-mainstream language uses. We contribute two newly annotated datasets that achieve higher inter-annotator agreement between human and language model (LLM) annotations compared to original datasets based on descriptive instructions. Our experiments show that LLMs can serve as effective alternatives when professional annotators are unavailable. Moreover, smaller models fine-tuned on multi-source LLM-annotated data outperform models trained on larger, single-source human-annotated datasets. These findings highlight the value of structured guidelines in reducing subjective variability, maintaining performance with limited data, and embracing language diversity. Content Warning: This article only analyzes offensive language for academic purposes. Discretion is advised.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。