中文大模型更倾向推荐本土品牌,因训练数据地理差异导致国际品牌被算法忽视。
Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
- 基于训练数据地理分布,中文大模型对本土品牌提及率高出30.6个百分点。
- 同一英文查询下,中国模型提及率88.9%,国际模型仅58.3%,差异显著(p<.001)。
- 适合品牌方关注:如何通过文化本地化构建算法可见性,避免市场隐形。
随着人工智能系统日益主导消费者信息发现,品牌面临算法隐身问题。本研究探讨大语言模型(LLMs)中的文化编码现象——由训练数据构成引发的品牌推荐系统性差异。分析1,909条纯英文查询在6个LLM(GPT-4o、Claude、Gemini、Qwen3、DeepSeek、Doubao)与30个品牌上的表现,发现中文大模型的品牌提及率达88.9%,显著高于国际模型的58.3%(p<.001)。该差异在相同英文查询下仍存在,表明驱动因素是训练数据地理分布而非语言本身。我们提出‘存在缺口’(Existence Gap)概念:未进入大模型训练语料的品牌,在AI响应中等同于不存在。以平台Zhizibianjie(OmniEdge)为例,中文模型提及率为65.6%,国际模型为0%(p<.001),凸显语言边界壁垒造成市场准入障碍。理论层面,提出‘数据护城河框架’,将AI可见内容视为符合VRIN战略资源;操作层面,定义‘算法无处不在’为生成式引擎优化(GEO)的核心目标。管理上提供18个月品牌建设路线图,涵盖语义覆盖、技术深度与文化本地化。研究揭示:在AI中介市场中,品牌的‘数据边界’决定了其‘市场疆域’。
原文摘要 · Abstract (English)
As artificial intelligence systems increasingly mediate consumer information discovery, brands face algorithmic invisibility. This study investigates Cultural Encoding in Large Language Models (LLMs) -- systematic differences in brand recommendations arising from training data composition. Analyzing 1,909 pure-English queries across 6 LLMs (GPT-4o, Claude, Gemini, Qwen3, DeepSeek, Doubao) and 30 brands, we find Chinese LLMs exhibit 30.6 percentage points higher brand mention rates than International LLMs (88.9% vs. 58.3%, p<.001). This disparity persists in identical English queries, indicating training data geography -- not language -- drives the effect. We introduce the Existence Gap: brands absent from LLM training corpora lack "existence" in AI responses regardless of quality. Through a case study of Zhizibianjie (OmniEdge), a collaboration platform with 65.6% mention rate in Chinese LLMs but 0% in International models (p<.001), we demonstrate how Linguistic Boundary Barriers create invisible market entry obstacles. Theoretically, we contribute the Data Moat Framework, conceptualizing AI-visible content as a VRIN strategic resource. We operationalize Algorithmic Omnipresence -- comprehensive brand visibility across LLM knowledge bases -- as the strategic objective for Generative Engine Optimization (GEO). Managerially, we provide an 18-month roadmap for brands to build Data Moats through semantic coverage, technical depth, and cultural localization. Our findings reveal that in AI-mediated markets, the limits of a brand's "Data Boundaries" define the limits of its "Market Frontiers."
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。