用大模型少样本学习,快速判断新闻和视频的政治立场。
"Whose Side Are You On?" Estimating Ideology of Political and News Content Using Large Language Models and Few-shot Demonstration Selection
- 通过挑选代表性样例,让大模型在上下文中学会识别立场。
- 在三个数据集上表现优于零样本和传统监督方法。
- 发现内容来源信息能显著影响模型判断,适合舆情分析者使用。
社交媒体的快速发展引发了人们对极端化、信息茧房和内容偏见的担忧。现有意识形态分类方法受限于人工标注成本高、需大规模标注数据,且难以适应不断变化的意识形态语境。本文探索利用大型语言模型(LLMs)通过上下文学习(ICL)对在线内容进行政治意识形态分类。我们在包含新闻文章和YouTube视频的三个数据集上,系统地进行了标签平衡式样例选择实验,结果表明该方法显著优于零样本和传统监督方法。此外,我们评估了元数据(如内容来源、描述)对意识形态分类的影响,并讨论其意义。最后,我们展示了为政治与非政治内容提供来源信息如何影响大模型的分类结果。
原文摘要 · Abstract (English)
The rapid growth of social media platforms has led to concerns about radicalization, filter bubbles, and content bias. Existing approaches to classifying ideology are limited in that they require extensive human effort, the labeling of large datasets, and are not able to adapt to evolving ideological contexts. This paper explores the potential of Large Language Models (LLMs) for classifying the political ideology of online content through in-context learning (ICL). Our extensive experiments involving demonstration selection in label-balanced fashion, conducted on three datasets comprising news articles and YouTube videos, reveal that our approach significantly outperforms zero-shot and traditional supervised methods. Additionally, we evaluate the influence of metadata (e.g., content source and descriptions) on ideological classification and discuss its implications. Finally, we show how providing the source for political and non-political content influences the LLM's classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。