分析140万条评论,发现政治立场显著影响用户对LGBTQ+新闻的评价
Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech
- 构建细粒度希望言论分类器,识别评论中的希望、负面、中性内容
- 3750条标注数据揭示:政治立场越右,越倾向负面评价LGBTQ+内容
- 零样本大模型更贴近自由派评价,提示算法偏见需警惕
本文有三项贡献。首先,基于来自美国主要有线新闻媒体3,161个YouTube新闻视频的1,419,047条评论,分析用户对LGBTQ+新闻内容的互动行为,特别构建细粒度希望言论分类器,以识别正面(希望言论)、负面、中性和无关内容。其次,联合一位专注LGBTQ+健康的公共卫生专家,开展具有均衡多元政治背景的标注研究,发布包含3,750个实例的标注数据集,附带详细标注者人口统计信息。最后,研究揭示:(1)标注者政治立场与其对边缘群体相关内容的评价强相关;(2)基于个体政治立场训练的模型在真实场景中存在显著分歧;(3)零样本大语言模型(LLMs)的判断更接近自由派标注者。
原文摘要 · Abstract (English)
This paper makes three contributions. First, via a substantial corpus of 1,419,047 comments posted on 3,161 YouTube news videos of major US cable news outlets, we analyze how users engage with LGBTQ+ news content. Our analyses focus both on positive and negative content. In particular, we construct a fine-grained hope speech classifier that detects positive (hope speech), negative, neutral, and irrelevant content. Second, in consultation with a public health expert specializing on LGBTQ+ health, we conduct an annotation study with a balanced and diverse political representation and release a dataset of 3,750 instances with fine-grained labels and detailed annotator demographic information. Finally, beyond providing a vital resource for the LGBTQ+ community, our annotation study and subsequent in-the-wild assessments reveal (1) strong association between rater political beliefs and how they rate content relevant to a marginalized community; (2) models trained on individual political beliefs exhibit considerable in-the-wild disagreement; and (3) zero-shot large language models (LLMs) align more with liberal raters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。