用委婉语和贬损语测试大模型事实核查中的政治偏见
PolBiX: Detecting LLMs' Political Bias in Fact-Checking through X-phemisms
- 通过替换词语的褒贬色彩构造语义等价句子对
- 发现判断性词汇比政治立场更影响模型判断结果
- 适合关注AI伦理与偏见检测的研究者阅读
大型语言模型在需要客观判断的应用中日益普及,但其政治偏见可能影响评估质量。已有研究发现多数模型存在左倾倾向,但其在事实核查等下游任务中的具体影响仍不明确。本研究通过在德语文本中使用委婉语(euphemisms)和贬损语(dysphemisms)构建语义等价的最小差异句对,系统检验模型在判断真假时的一致性。我们评估了六种主流LLM,结果表明:相较于政治立场,判断性词汇的存在显著影响模型对真实性的判断。少数模型表现出政治偏见,但即使在提示中强调客观性,也无法有效缓解该问题。注意:本文包含可能令人不适的内容。
原文摘要 · Abstract (English)
Large Language Models are increasingly used in applications requiring objective assessment, which could be compromised by political bias. Many studies found preferences for left-leaning positions in LLMs, but downstream effects on tasks like fact-checking remain underexplored. In this study, we systematically investigate political bias through exchanging words with euphemisms or dysphemisms in German claims. We construct minimal pairs of factually equivalent claims that differ in political connotation, to assess the consistency of LLMs in classifying them as true or false. We evaluate six LLMs and find that, more than political leaning, the presence of judgmental words significantly influences truthfulness assessment. While a few models show tendencies of political bias, this is not mitigated by explicitly calling for objectivism in prompts. Warning: This paper contains content that may be offensive or upsetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。