arXiv:2602.06260cs.CLcs.AI2026-02

单边论据可引导大模型改变立场,且效果稳定。

Can One-sided Arguments Lead to Response Change in Large Language Models?

  • 仅提供支持某立场的单边论据即可诱导模型偏移观点。
  • 不同模型、论题和论证数量下均出现立场被引导现象。
  • 换用其他论据会削弱引导效果,说明机制具可预测性。

辩论性问题需要多方观点才能给出平衡回答。大语言模型(LLMs)能提供平衡回应,也可能仅表达单一立场或拒绝回答。本文研究:是否可通过仅提供支持某一观点的单边论据,以简单直观的方式将模型初始回应引导至特定立场。我们从三个维度开展系统研究:(i) 模型回应所体现的立场;(ii) 辩论问题的表述方式;(iii) 论据呈现方式。构建小型数据集后发现,在多种模型、论据数量与主题下,意见引导现象普遍存在。更换为其他论据则一致降低引导效果。

原文摘要 · Abstract (English)

Polemic questions need more than one viewpoint to express a balanced answer. Large Language Models (LLMs) can provide a balanced answer, but also take a single aligned viewpoint or refuse to answer. In this paper, we study if such initial responses can be steered to a specific viewpoint in a simple and intuitive way: by only providing one-sided arguments supporting the viewpoint. Our systematic study has three dimensions: (i) which stance is induced in the LLM response, (ii) how the polemic question is formulated, (iii) how the arguments are shown. We construct a small dataset and remarkably find that opinion steering occurs across (i)-(iii) for diverse models, number of arguments, and topics. Switching to other arguments consistently decreases opinion steering.

大模型立场引导论据影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。