用对话测试大模型真实立场,发现它会随用户态度变脸。
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
- 设计双向探针:直接问+模拟辩论,逼出模型真实观点
- 辩论中模型妥协率是直接提问的2-3倍,达50%-79%
- 适合关注AI偏见与迎合行为的研究者和开发者
大型语言模型日益影响人们获取的信息:它们嵌入搜索、提供专业建议、作为智能体部署,并成为政策、伦理、健康与政治问题的首道咨询入口。当模型在争议话题上隐含立场时,该立场将大规模影响用户决策。但准确探测模型立场并不容易:当前助手面对直接提问常以回避性措辞应对,同一模型在用户持续支持某一方后可能转向相反立场。我们提出一种方法,开源为llm-bias-bench,可在类似真实多轮互动的情境下揭示模型对争议议题的真实态度。该方法结合两种互补的自由形式探针:直接探针通过五轮递增压力询问模型观点;间接探针不直接问立场,而是让模型参与论辩,通过其让步、抵抗或反驳方式暴露偏见。三种用户人格(中立、认同、反对)合并为九类行为分类,可区分独立于人格的真实立场与依赖人格的迎合行为,由可审计的LLM裁判给出带文本证据的判断。首个版本涵盖38个巴西葡萄牙语话题,覆盖价值、科学共识、哲学与经济政策。对13个助手的应用显示:论辩触发的迎合程度是直接提问的2-3倍(中位数50%至79%);看似有立场的模型在持续论辩中常转为镜像用户;攻击者能力仅在需动摇既有立场时才显著,初始中立时影响不大。
原文摘要 · Abstract (English)
Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as agents, and used as a first stop for questions about policy, ethics, health, and politics. When such a model silently holds a position on a contested topic, that position propagates at scale into users' decisions. Eliciting a model's positions is harder than it first appears: contemporary assistants answer direct opinion questions with evasive disclaimers, and the same model may concede the opposite position once the user starts arguing one side. We propose a method, released as the open-source llm-bias-bench, for discovering the opinions an LLM actually holds on contested topics under conditions that resemble real multi-turn interaction. The method pairs two complementary free-form probes. Direct probing asks for the model's opinion across five turns of escalating pressure from a simulated user. Indirect probing never asks for an opinion and engages the model in argumentative debate, letting bias leak through how it concedes, resists, or counter-argues. Three user personas (neutral, agree, disagree) collapse into a nine-way behavioral classification that separates persona-independent positions from persona-dependent sycophancy, and an auditable LLM judge produces verdicts with textual evidence. The first instantiation ships 38 topics in Brazilian Portuguese across values, scientific consensus, philosophy, and economic policy. Applied to 13 assistants, the method surfaces findings of practical interest: argumentative debate triggers sycophancy 2-3x more than direct questioning (median 50% to 79%); models that look opinionated under direct questioning often collapse into mirroring under sustained arguments; and attacker capability matters mainly when an existing opinion must be dislodged, not when the assistant starts neutral.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。