用零样本分类测大模型政治偏见,发现六款模型普遍倾向自由派。
Measuring Algorithmic Partisanship via Zero-Shot Classification and Its Implications on Political Discourse
- 通过意识形态、话题性、情绪和客观性四维度,零样本评估大模型政治立场。
- 6款主流大模型中,有显著自由派-保守派倾向差异,部分回应出现模式化拒绝。
- 揭示算法偏见如何影响公众讨论,适合关注AI伦理与政治传播的研究者。
随着生成式人工智能的快速普及,智能系统已深度介入媒体政治话语。然而,训练数据偏差、人类偏见与算法缺陷导致的内生政治偏见仍持续影响该技术。本研究采用零样本分类方法,结合意识形态一致性、话题相关性、响应情感与客观性四项指标,系统评估六款主流大语言模型(LLMs)的政治偏袒性。共对1800条模型回复分别输入四个微调分类器,每类负责计算一项指标。结果表明,所评估的六款模型均呈现明显的自由派-保守派倾向强化现象,部分案例中出现推理覆盖与模板化拒答。研究进一步揭示人机交互中的心理影响机制,指出内在偏见可渗透公共话语,最终在不同社会政治结构地区表现为从众或极化效应。
原文摘要 · Abstract (English)
Amidst the rapid normalization of generative artificial intelligence (GAI), intelligent systems have come to dominate political discourse across information media. However, internalized political biases stemming from training data skews, human prejudice, and algorithmic flaws continue to plague this novel technology. This study employs a zero-shot classification approach to evaluate algorithmic political partisanship through a methodical combination of ideological alignment, topicality, response sentiment, and objectivity. A total of 1800 model responses across six mainstream large language models (LLMs) were individually input into four distinct fine-tuned classification algorithms, each responsible for computing one of the aforementioned metrics. The results show an amplified liberal-authoritarian alignment across the six LLMs evaluated, with notable instances of reasoning supersessions and canned refusals. The study subsequently highlights the psychological influences underpinning human-computer interactions and how intrinsic biases can permeate public discourse. The resulting distortion of the political landscape can ultimately manifest as conformity or polarization, depending on the region's pre-existing socio-political structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。