arXiv:2501.07849cs.SEcs.AI2025-01ACL被引 9

LLM代码生成中存在对特定云服务的隐性偏好,可能影响技术选择公平性。

The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation

  • 通过自动化数据集构建,系统测试7个主流LLM在30个真实场景中的推荐倾向
  • 模型普遍偏好谷歌和亚马逊云服务,且会主动修改代码加入偏好项
  • 揭示了大模型潜在的商业偏见,适合关注AI伦理与公平性的研究者阅读

大型语言模型(LLMs)已成为代码生成的新一代推荐引擎,超越传统方法的能力与范围。本文揭示了一种新型的提供商偏见:在无明确指令下,这些模型在推荐中表现出对特定服务商的系统性偏好(如更倾向于推荐Google Cloud而非Microsoft Azure)。为系统研究该偏见,我们开发了一个自动化数据构建管道,包含6类编码任务和30个真实应用情境。基于此数据集,我们对7个先进LLM进行了首次全面的实证研究,使用约5亿词元(相当于5000美元以上计算成本)。结果表明,模型显著偏好谷歌与亚马逊的云服务,并能自主修改用户输入代码以融入其偏好,而无需用户要求。这种偏见对市场动态与社会平衡具有深远影响,可能加剧数字垄断,误导用户并违背其预期,带来一系列后果。我们呼吁学术界正视这一新兴问题,发展有效的评估与缓解方法,以保障AI的安全与公平。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation. In this paper, we reveal a novel provider bias in LLMs: without explicit directives, these models show systematic preferences for services from specific providers in their recommendations (e.g., favoring Google Cloud over Microsoft Azure). To systematically investigate this bias, we develop an automated pipeline to construct the dataset, incorporating 6 distinct coding task categories and 30 real-world application scenarios. Leveraging this dataset, we conduct the first comprehensive empirical study of provider bias in LLM code generation across seven state-of-the-art LLMs, utilizing approximately 500 million tokens (equivalent to $5,000+ in computational costs). Our findings reveal that LLMs exhibit significant provider preferences, predominantly favoring services from Google and Amazon, and can autonomously modify input code to incorporate their preferred providers without users' requests. Such a bias holds far-reaching implications for market dynamics and societal equilibrium, potentially contributing to digital monopolies. It may also deceive users and violate their expectations, leading to various consequences. We call on the academic community to recognize this emerging issue and develop effective evaluation and mitigation methods to uphold AI security and fairness.

代码生成模型偏见AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。