用推理时策略提升大模型政治中立性,防止恶意提示注入偏见。
Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

- 通过思维链和直接偏好优化,在推理阶段缓解偏见注入。
- 在立法视频摘要任务中,中立性评分从2.14提升至4.56(满分5分)。
- 适合关注AI安全、内容生成可信度的研究者与应用开发者。
随着大语言模型(LLMs)成为信息检索与摘要任务的核心工具,确保其始终保持非党派性并抵御政治偏见,是实现更安全、更可信人工智能的关键一步。当前的对齐范式(如基于人类反馈的强化学习)虽能引导模型遵循安全指令,但可能被对抗性提示注入利用,生成不安全内容。尤其政治偏见未被现代对齐技术明确视为有害内容。为应对这一漏洞,本文提出基于思维链(CoT)提示与直接偏好优化(DPO)的缓解策略。我们使用公开的立法视频数据集,生成摘要并注入偏见,通过四轴政治摘要评估体系进行评测。结果表明,所提出的递归自我修正方法使模型在政治中立性李克特量表上的平均得分从2.14提升至4.56,验证了推理时缓解政治偏见的有效性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI). Current model alignment paradigms, such as reinforcement learning from human feedback (RLHF), make LLMs follow overarching safety instructions. However, this instruction tuning can be exploited via adversarial prompt injection and be used to generate unsafe content. In particular, political bias has not been specifically targeted by modern alignment techniques as harmful and biased content. To address this vulnerability of LLMs, we propose mitigation strategies using Chain of Thought (CoT) prompting and Direct Preference Optimization (DPO). Using a public dataset of legislative videos, we generate summaries using LLMs, inject bias via adversarial prompting and evaluate their performance on a four axis scale designed for political summarization. In this paper, we present different methods to shield LLMs against the injection of political bias. Our results demonstrate that the proposed Recursive Self-Correction approach raises model performance from a Political Neutrality Likert scale baseline of 2.14 to 4.56, averaged across all models, demonstrating effective inference-time mitigation of political bias in LLM-generated summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。