用大模型解析政治言论中的对象与立场,实现细粒度分析。
Large Language Models Unpack Complex Political Opinions through Target-Stance Extraction
- 通过目标-立场抽取任务,让大模型识别政治言论中的讨论对象和态度。
- 最佳模型表现接近专业人工标注,对争议性强的帖子也保持稳定。
- 适合做政治话语分析、社会科学研究的学者与机构使用。
政治极化源于政策、人物和议题信念之间的复杂互动。然而,多数计算分析仅将话语简化为粗略的党派标签,忽略了这些信念如何相互作用。这在在线政治对话中尤为明显,其内容往往细腻多元,难以自动识别讨论对象及对应立场。本研究探讨大语言模型(LLMs)能否通过目标-立场抽取(TSE)任务解决此问题,该任务结合目标识别与立场检测,实现更精细的政治观点分析。我们构建了一个包含1,084篇r/NeutralPolitics Reddit帖子的数据集,涵盖138个不同政治目标,并评估多种专有及开源LLMs,采用零样本、少样本和上下文增强提示策略。结果表明,最优模型表现与高度训练的人类标注者相当,且在标注一致性低的挑战性帖子上仍具鲁棒性。研究证明,大模型可在极少监督下提取复杂政治意见,为计算社会科学与政治文本分析提供可扩展工具。
原文摘要 · Abstract (English)
Political polarization emerges from a complex interplay of beliefs about policies, figures, and issues. However, most computational analyses reduce discourse to coarse partisan labels, overlooking how these beliefs interact. This is especially evident in online political conversations, which are often nuanced and cover a wide range of subjects, making it difficult to automatically identify the target of discussion and the opinion expressed toward them. In this study, we investigate whether Large Language Models (LLMs) can address this challenge through Target-Stance Extraction (TSE), a recent natural language processing task that combines target identification and stance detection, enabling more granular analysis of political opinions. For this, we construct a dataset of 1,084 Reddit posts from r/NeutralPolitics, covering 138 distinct political targets and evaluate a range of proprietary and open-source LLMs using zero-shot, few-shot, and context-augmented prompting strategies. Our results show that the best models perform comparably to highly trained human annotators and remain robust on challenging posts with low inter-annotator agreement. These findings demonstrate that LLMs can extract complex political opinions with minimal supervision, offering a scalable tool for computational social science and political text analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。