对比大模型在政治文本标注中的表现,发现开源模型兼具高效与可复现性。
Benchmarking LLMs in Political Content Text-Annotation: Proof-of-Concept with Toxicity and Incivility Data
- 用三百万条社交互动数据构建真实标注基准,测试多模型零样本分类能力。
- GPT-4o和Nous Hermes 2 Mixtral在毒性与不文明内容识别上表现最佳,准确率领先。
- 小参数开源模型如Nous Hermes 2也能高精度完成任务,适合资源受限场景。
本文评估了OpenAI的GPT系列及多个开源大模型在政治类文本标注任务中的表现。基于包含三百万条数字互动的新型抗议事件数据集,构建了由人工标注的黄金标准标签集,涵盖社交媒体上的毒性与不文明内容。测试中,Google Perspective API(采用宽松阈值)、GPT-4o以及Nous Hermes 2 Mixtral在零样本分类任务中表现最优。此外,参数量更小的Nous Hermes 2与Mistral OpenOrca也展现出高水平性能,是兼顾效果、成本与计算效率的理想选择。附加实验表明,尽管GPT系列在响应速度和可靠性方面优异,但仅开源模型能保证标注过程的完全可复现性。
原文摘要 · Abstract (English)
This article benchmarked the ability of OpenAI's GPTs and a number of open-source LLMs to perform annotation tasks on political content. We used a novel protest event dataset comprising more than three million digital interactions and created a gold standard that includes ground-truth labels annotated by human coders about toxicity and incivility on social media. We included in our benchmark Google's Perspective algorithm, which, along with GPTs, was employed throughout their respective APIs while the open-source LLMs were deployed locally. The findings show that Perspective API using a laxer threshold, GPT-4o, and Nous Hermes 2 Mixtral outperform other LLM's zero-shot classification annotations. In addition, Nous Hermes 2 and Mistral OpenOrca, with a smaller number of parameters, are able to perform the task with high performance, being attractive options that could offer good trade-offs between performance, implementing costs and computing time. Ancillary findings using experiments setting different temperature levels show that although GPTs tend to show not only excellent computing time but also overall good levels of reliability, only open-source LLMs ensure full reproducibility in the annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。