arXiv:2411.11081cs.CL2024-11NAACL被引 31

用大模型自动生成媒体偏见数据集,成本降了但效果有局限。

The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection

  • 用大模型批量生成4.8万条媒体偏见标注数据
  • 自动生成数据训练的模型比所有大模型标注者高5-9%准确率
  • 适合想低成本构建偏见检测模型的研究者

高昂的人工标注成本阻碍了高质量文本分类数据集的构建。近期研究提出使用大语言模型(LLMs)自动化标注,降低费用同时保持数据质量。尽管在仇恨言论检测和政治话语分析中表现良好,但其在复杂媒体偏见检测任务中的可行性仍待验证。本研究构建了annolexical——首个大规模媒体偏见分类数据集,包含超过48000条合成标注样本。基于该数据集微调的分类器,在两个基准数据集(BABE和BASIL)上,相比所有标注用的LLMs,Matthews相关系数(MCC)提升5-9个百分点,且接近或优于人类标注数据训练的模型。该方法显著降低了媒体偏见领域数据集构建与模型开发的成本。然而行为压力测试也揭示了当前方法的部分局限与权衡。

原文摘要 · Abstract (English)

High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality. LLMs have shown promising results in annotating downstream tasks like hate speech detection and political framing. Building on the success in these areas, this study investigates whether LLMs are viable for annotating the complex task of media bias detection and whether a downstream media bias classifier can be trained on such data. We create annolexical, the first large-scale dataset for media bias classification with over 48000 synthetically annotated examples. Our classifier, fine-tuned on this dataset, surpasses all of the annotator LLMs by 5-9 percent in Matthews Correlation Coefficient (MCC) and performs close to or outperforms the model trained on human-labeled data when evaluated on two media bias benchmark datasets (BABE and BASIL). This study demonstrates how our approach significantly reduces the cost of dataset creation in the media bias domain and, by extension, the development of classifiers, while our subsequent behavioral stress-testing reveals some of its current limitations and trade-offs.

媒体偏见大模型标注数据集构建自动化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。