测试大模型在不同视角设定下的仇恨言论标注表现,发现其倾向聚合观点而非个性化判断。
The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum
- 用预设角色身份引导大模型进行内容标注,探索视角多样性影响
- 弱视角下模型表现优于强视角和真人标注,体现观点聚合倾向
- 在个性化数据集上接近人类表现,但未超越,适合需统一标准的场景
本研究探讨大型语言模型(LLMs)在强到弱数据透视谱系中,基于预设标注者角色身份对仇恨言论与攻击性内容进行标注的能力。我们评估了模型生成标注与现有视角建模技术的表现。结果表明,LLMs会选择性使用角色中的社会人口属性。我们识别出典型标注者,其角色特征与原始人类标注者存在不同程度的契合度。在数据透视范式下,不依赖标注者信息的建模方法在弱数据透视条件下表现优于强数据透视及人类标注,表明大模型生成的观点倾向于聚合,即使在主观提示下亦如此。然而,在针对强透视设计的个性化数据集上,大模型标注性能接近但未超过人类标注者。
原文摘要 · Abstract (English)
In this work, we explore the capability of Large Language Models (LLMs) to annotate hate speech and abusiveness while considering predefined annotator personas within the strong-to-weak data perspectivism spectra. We evaluated LLM-generated annotations against existing annotator modeling techniques for perspective modeling. Our findings show that LLMs selectively use demographic attributes from the personas. We identified prototypical annotators, with persona features that show varying degrees of alignment with the original human annotators. Within the data perspectivism paradigm, annotator modeling techniques that do not explicitly rely on annotator information performed better under weak data perspectivism compared to both strong data perspectivism and human annotations, suggesting LLM-generated views tend towards aggregation despite subjective prompting. However, for more personalized datasets tailored to strong perspectivism, the performance of LLM annotator modeling approached, but did not exceed, human annotators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。