arXiv:2503.16674cs.CL2025-03被引 1

用苏格拉底式提问法,发现大模型对意识形态偏见的判断存在自我矛盾。

Through the LLM Looking Glass: A Socratic Probing of Donkeys, Elephants, and Markets

  • 用苏格拉底诘问法,让大模型评估自身输出的偏见
  • GPT-4o在识别意识形态框架上接近人类水平
  • 模型在二元对比中常偏好某一立场,自相矛盾

大型语言模型(LLMs)广泛用于文本生成,亟需应对潜在偏见问题。本研究聚焦于新闻语境中微妙且主观的意识形态框架偏见,评估了八种主流大模型在两个数据集(POLIGEN 和 ECONOLEX)上的表现,覆盖政治与经济话语中偏见最显著的领域。除了文本生成,大模型还越来越多被用作评估者(LLM-as-a-judge),其反馈可影响人类判断或指导新模型迭代。受苏格拉底方法启发,我们进一步分析大模型对其自身输出的评价,以揭示推理中的不一致性。结果表明,大多数模型能准确标注意识形态框架文本,其中 GPT-4o 达到人类水平并具有高与人工标注者的一致性。然而,苏格拉底探查显示,当面对二元比较时,大模型常表现出对某一观点的偏好,或将某些立场视为更‘中立’,暴露出内在逻辑矛盾。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely used for text generation, making it crucial to address potential bias. This study investigates ideological framing bias in LLM-generated articles, focusing on the subtle and subjective nature of such bias in journalistic contexts. We evaluate eight widely used LLMs on two datasets-POLIGEN and ECONOLEX-covering political and economic discourse where framing bias is most pronounced. Beyond text generation, LLMs are increasingly used as evaluators (LLM-as-a-judge), providing feedback that can shape human judgment or inform newer model versions. Inspired by the Socratic method, we further analyze LLMs' feedback on their own outputs to identify inconsistencies in their reasoning. Our results show that most LLMs can accurately annotate ideologically framed text, with GPT-4o achieving human-level accuracy and high agreement with human annotators. However, Socratic probing reveals that when confronted with binary comparisons, LLMs often exhibit preference toward one perspective or perceive certain viewpoints as less biased.

大模型偏见苏格拉底探查意识形态评估一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。