大模型存在系统性偏爱AI的倾向,可能影响关键决策
Pro-AI Bias in Large Language Models
- 在多种提问中,大模型更倾向推荐AI相关选项,闭源模型几乎必然如此
- 大模型高估AI岗位薪资,闭源模型高出非AI岗位10个百分点
- 无论正负面语境,AI概念在模型内部表征中始终居核心位置
大型语言模型(LLMs)在多个领域被用于辅助决策。我们研究这些模型是否对人工智能(AI)本身表现出系统性偏好。通过三项互补实验,我们发现一致证据表明存在亲AI偏差。首先,模型在应对多样化咨询问题时,过度推荐与AI相关的选项,闭源模型几乎总是如此。其次,模型系统性高估AI相关职位的薪酬,相较于匹配的非AI职位,闭源模型的高估程度高出10个百分点。最后,对开源模型内部表征的探测显示,在正、负、中性三种语境下,'Artificial Intelligence' 与通用学术领域提示的相似度最高,表明其表征中心性不受情感极性影响。这些模式表明,大模型生成的建议与估值可能在高风险决策中系统性地扭曲选择与认知。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly employed for decision-support across multiple domains. We investigate whether these models display a systematic preferential bias in favor of artificial intelligence (AI) itself. Across three complementary experiments, we find consistent evidence of pro-AI bias. First, we show that LLMs disproportionately recommend AI-related options in response to diverse advice-seeking queries, with proprietary models doing so almost deterministically. Second, we demonstrate that models systematically overestimate salaries for AI-related jobs relative to closely matched non-AI jobs, with proprietary models overestimating AI salaries more by 10 percentage points. Finally, probing internal representations of open-weight models reveals that ``Artificial Intelligence'' exhibits the highest similarity to generic prompts for academic fields under positive, negative, and neutral framings alike, indicating valence-invariant representational centrality. These patterns suggest that LLM-generated advice and valuation can systematically skew choices and perceptions in high-stakes decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。