分析大模型如何跨语言跨市场获取品牌信息,发现来源高度依赖第三方网站。
How Large Language Models Source Brand Reputation Across Languages and Markets
- 追踪大模型回答中的引用链接,分析其信息来源分布。
- 85.7%的品牌信息来自非官方第三方网站,仅14.3%来自品牌自建网站。
- 维基百科几乎主导所有语言的引用,但个别市场如波兰受视频平台影响显著。
当大型语言模型(LLM)回答关于公司的问题时,其答案基于检索到的网页源,而这些来源决定了模型的内容。现有对AI品牌可见性的研究多关注回答文本本身,本文则回溯一步,聚焦于引用来源。我们整合了三个Rankfor.AI数据集,覆盖128个品牌、12个母市场和13种语言,共分析167,551条以URL为依据的引用(总计189,974条归因记录)。按域名和来源类型对每条引用进行分类,测量AI在不同语言和市场中获取品牌信息的来源分布。四个规律成立:第一,大模型几乎全部依赖第三方来源,85.7%的引用指向品牌不拥有的网站,仅14.3%为品牌自有;第二,来源集中且呈长尾分布,约18%的域名贡献了80%的引用,符合齐普夫定律(alpha = 0.86,R² = 0.983);第三,维基百科在11种语言中是最常被引用的域名,仅立陶宛例外,当地商业日报vz.lt略胜一筹(占比4.38%);第四,市场差异在边缘显现:对46个波兰本土品牌而言,最常被引用的是YouTube,四个人力资源与职业类门户贡献637条引用,远超波兰维基百科的297条,约为其两倍。
原文摘要 · Abstract (English)
When a large language model (LLM) answers a question about a company, it grounds the answer in retrieved web sources, and those sources decide what the model says. Most analysis of AI brand visibility looks at the answer text. This study looks one step earlier, at the citations. We merge three Rankfor.AI datasets covering 128 brands across 12 home markets and 13 languages, and analyse 167,551 URL-grounded citations (189,974 total attribution rows). We classify each citation by domain and source type and measure where AI gets its brand information, by language and by market. Four patterns hold. First, AI grounds brand answers overwhelmingly in third-party sources: 85.7% of citations point to sites the brand does not own, against 14.3% owned. Second, the source base is concentrated and long-tailed: 80% of citations come from about 18% of domains, fitting a Zipf law (alpha = 0.86, R^2 = 0.983). Third, one reference site dominates almost everywhere: Wikipedia is the most-cited domain in 11 of 12 languages, the exception being Lithuanian, where the business daily vz.lt edges it (4.38%). Fourth, the source mix is market-specific at the margin: for 46 Polish national brands the most-cited domain is YouTube, and four HR and careers portals supply 637 citations against 297 for Polish Wikipedia, about twice as many.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。