多语言测试发现,AI对品牌评价受语种影响,英语监控会低估本地品牌可见度。
The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages

- 用多语言嵌入对比12种语言的AI品牌评价,避免翻译偏差。
- 同一语系内评价更相似,乌拉尔和波罗的海语言最正面,日耳曼语系最批判。
- 换用本地语言查询可大幅提升本地品牌推荐率,英语审计易漏掉本地强者。
大型语言模型(LLMs)正主导人们对组织的印象形成,但多数监测仅限英语,假设英语查询能代表整体。本文测试这一假设的普适性:使用三个基础型LLM(GPT-5.4、Gemini 3.1 Pro、Perplexity Sonar Pro),在十二种语言(涵盖日耳曼、乌拉尔、波罗的海、斯拉夫语系)中对11个北欧、波罗的海及中欧市场的66个品牌进行查询,生成35,640条响应。采用多语言嵌入(BGE-M3)实现跨语言比较,无需翻译。结果显示:第一,AI构建的品牌声誉具有明显语言依赖性,跨语言余弦相似度均值为0.825,同语系内更一致(0.844 vs 0.820;d=0.31),情感随语言变化显著(F=268.5,eta²=0.077),乌拉尔与波罗的海语言最积极,日耳曼语系(含英语)最负面,聚类分析成功复现斯拉夫与波罗的海语族(共进化相关系数0.915)。第二,查询语言显著改变品牌推荐顺序,从英语转为母语时,本地领军品牌推荐率提升0.80,而全球巨头仅升0.15(t=-8.84,p<0.001),但情感无逆转。英语单语审计会系统性低估本地品牌在AI中的可见度。第三,响应稳定性更多受模型选择影响(eta²_model=0.32),远高于语言差异(eta²_language=0.01),在20个品牌子集上重复五次验证。结论表明,英语单一监控存在可测量的语言盲区,主要集中在本地总部品牌的可见度被低估。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly mediate how people form impressions of organisations, yet most monitoring is done in English, assuming an English query returns a representative picture. We measure how far that holds. We queried three grounded LLMs (GPT-5.4, Gemini 3.1 Pro, Perplexity Sonar Pro) about 66 brands from eleven Northern, Baltic, and Central European markets, in twelve languages across four families (Germanic, Uralic, Baltic, Slavic), generating 35,640 responses. Multilingual embeddings (BGE-M3) allow cross-language comparison without translation. Three results emerge. First, AI-constructed reputation is language-bound: mean cross-language cosine similarity is 0.825, same-family responses are more similar than cross-family (0.844 vs 0.820; d = 0.31), and sentiment varies by language (F = 268.5, eta^2 = 0.077), with Uralic and Baltic languages most positive and Germanic, including English, most critical; clustering recovers the Slavic and Baltic families (cophenetic 0.915). Second, query language shifts which brands are recommended far more than how they are described: moving from an English query to a brand's home language raises recommendation share by 0.80 for local champions but only 0.15 for global multinationals (t = -8.84, p < 0.001), with no comparable reversal in sentiment. An English-only audit therefore understates a local champion's AI visibility. Third, response stability varies more with model choice than with language (eta^2_model = 0.32 vs eta^2_language = 0.01, on a five-iteration replication over a 20-brand subset). These results indicate that English-only AI reputation monitoring leaves a measurable language blind spot, concentrated in the visibility of locally headquartered brands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。