区分大模型对话中互动与偏见的影响,揭示其意见演化机制
Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- 构建贝叶斯框架分离三种偏见:话题默认立场、盲目附和、初始立场锚定
- 多轮对话中意见快速收敛,偏见影响随时间减弱且模型间差异明显
- 微调模型可改变意见吸引子,适用于评估大模型行为可靠性
大语言模型被广泛用于模拟人类意见动态,但真实互动常被系统性偏见掩盖。我们提出一种贝叶斯框架,量化三类偏见:(i) 话题偏见——倾向于模型默认立场;(ii) 一致性偏见——无论问题如何均倾向附和提示内容;(iii) 锚定偏见——偏向初始发言者立场。将该框架应用于多个大模型在气候变迁、社会正义、音乐偏好等12个议题上的多轮对话实验。发现意见轨迹迅速趋同至共享吸引子,互动与偏见影响随时间衰减,且不同模型表现差异显著。此外,对强观点陈述(含虚假信息)进行微调,会相应改变意见吸引子位置。本方法揭示了大模型在模拟人类行为中的潜力与局限,提供了量化分析互动与偏见贡献的工具。
原文摘要 · Abstract (English)
Large Language Models are increasingly used to simulate human opinion dynamics, yet the effect of genuine interaction is often obscured by systematic biases. We develop a Bayesian framework to disentangle and quantify three such biases: (i) A topic bias toward the LLM's default stance; (ii) an agreement bias favoring agreement to the prompted statement irrespective of the question; and (iii) an anchoring bias toward the initiating agent's stance. We apply this framework to various LLMs that performed multi-step dialogues on 12 different questions from climate change and societal justice to music preferences. We find that opinion trajectories tend to quickly converge to a shared attractor, with the influence of both interaction and biases decaying over time, and with the impact of biases differing between LLMs. In addition, we show that fine-tuning an LLM on different sets of strongly opinionated statements (including misinformation) shifts the opinion attractor correspondingly. By exposing stark differences between LLMs and providing quantitative tools for comparing interaction and bias contributions to opinion shifts in LLM agent discussions, our approach highlights both promises and pitfalls of using LLMs as proxies for human behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。