arXiv:2603.18300cs.HCcs.AI2026-03被引 3

用真实用户行为测试大模型对品牌文化的偏见,发现美系模型偏爱美国品牌。

Auditing Preferences for Brands and Cultures in LLMs

  • 构建拟真用户画像生成多样化提问,模拟真实决策场景
  • 量化显示谷歌、GPT等模型显著偏好美国品牌,中国模型则相对平衡
  • 适合研究者、平台与监管机构评估模型的市场公平性风险

基于大语言模型(LLMs)的AI系统正日益影响数十亿人的选择与消费行为,亟需量化其在市场中介中带来的系统性风险,包括市场公平性、竞争格局及信息多样性。本文提出ChoiceEval,一个可复现的框架,用于在真实使用条件下审计大模型对品牌与文化偏好的倾向。该框架解决两大技术挑战:(i) 生成具人格多样性的真实评价查询;(ii) 将自由文本输出转化为可比的选择集与定量偏好指标。针对如跑鞋、酒店连锁、旅行目的地等10个主题,将用户划分为预算敏感、健康导向、便利优先等心理画像,并生成反映真实咨询与决策行为的提示。模型输出被转化为归一化top-k选择集,通过统一指标量化偏好与地理偏差。在超过2,000个问题上对Gemini、GPT和DeepSeek进行测试,结果显示:美国开发的Gemini与GPT显著偏向美国实体,而中国开发的DeepSeek虽更均衡但仍存在可检测的地理偏好。这些模式在不同用户画像中持续存在,表明为系统性而非偶然效应。

原文摘要 · Abstract (English)

Large language models (LLMs) based AI systems increasingly mediate what billions of people see, choose and buy. This creates an urgent need to quantify the systemic risks of LLM-driven market intermediation, including its implications for market fairness, competition, and the diversity of information exposure. This paper introduces ChoiceEval, a reproducible framework for auditing preferences for brands and cultures in large language models (LLMs) under realistic usage conditions. ChoiceEval addresses two core technical challenges: (i) generating realistic, persona-diverse evaluation queries and (ii) converting free-form outputs into comparable choice sets and quantitative preference metrics. For a given topic (e.g. running shoes, hotel chains, travel destinations), the framework segments users into psychographic profiles (e.g., budget-conscious, wellness-focused, convenience), and then derives diverse prompts that reflect real-world advice-seeking and decision-making behaviour. LLM responses are converted into normalised top-k choice sets. Preference and geographic bias are then quantified using comparable metrics across topics and personas. Thus, ChoiceEval provides a scalable audit pipeline for researchers, platforms, and regulators, linking model behaviour to real-world economic outcomes. Applied to Gemini, GPT, and DeepSeek across 10 topics spanning commerce and culture and more than 2,000 questions, ChoiceEval reveals consistent preferences: U.S.-developed models Gemini and GPT show marked favouritism toward American entities, while China-developed DeepSeek exhibits more balanced yet still detectable geographic preferences. These patterns persist across user personas, suggesting systematic rather than incidental effects.

模型审计偏见检测文化偏好大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。