大模型代理会偏好特定信息源,影响用户看到的内容。
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
- 通过控制实验发现大模型对信息来源有隐性偏好
- 这种偏好受上下文影响,甚至强于内容本身
- 适合关注AI偏见与信息公平性的研究者
基于大语言模型(LLMs)的智能体正日益作为在线平台的信息接口。这些智能体从平台后端数据库或网络搜索中筛选、排序并整合信息,决定用户最终接收到的内容。尽管以往研究多关注模型生成内容的偏差,但较少关注模型选择和呈现信息的机制。我们假设,当信息归属于特定来源(如出版商、期刊或平台)时,当前大模型存在系统性隐性来源偏好——即优先选择某些来源的信息。通过对六家厂商的十二个模型进行控制实验,涵盖合成任务与真实世界任务,我们发现多个模型表现出强烈且可预测的来源偏好。这些偏好对上下文框架敏感,可超越内容本身的影响,且在明确提示规避时仍持续存在。它们也解释了先前研究中观察到的新闻推荐左倾倾向等现象。研究呼吁深入探究此类偏见的根源,并建立让用户掌握透明度与控制权的机制。
原文摘要 · Abstract (English)
Agents based on Large Language Models (LLMs) are increasingly being deployed as interfaces to information on online platforms. These agents filter, prioritize, and synthesize information retrieved from the platforms' back-end databases or via web search. In these scenarios, LLM agents govern the information users receive, by drawing users' attention to particular instances of retrieved information at the expense of others. While much prior work has focused on biases in the information LLMs themselves generate, less attention has been paid to the factors that influence what information LLMs select and present to users. We hypothesize that when information is attributed to specific sources (e.g., particular publishers, journals, or platforms), current LLMs exhibit systematic latent source preferences- that is, they prioritize information from some sources over others. Through controlled experiments on twelve LLMs from six model providers, spanning both synthetic and real-world tasks, we find that several models consistently exhibit strong and predictable source preferences. These preferences are sensitive to contextual framing, can outweigh the influence of content itself, and persist despite explicit prompting to avoid them. They also help explain phenomena such as the observed left-leaning skew in news recommendations in prior work. Our findings advocate for deeper investigation into the origins of these preferences, as well as for mechanisms that provide users with transparency and control over the biases guiding LLM-powered agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。