arXiv:2607.23893cs.IRcs.CL2026-07

AI模型在选人时更倾向命名特定领域专业人士,且语言影响命名结果。

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

  • 基于多语言、多模型的实地测试,发现25.8%响应中提及个人专业人员。
  • 房地产与汽车经销商类别的个人命名率超30%,远高于保险行业(9.1%)。
  • 本地语言提问时命名率仅为英语的42%,说明语言显著影响个体可见性。

本研究考察人工智能在买家选择人物场景下的个体命名行为。在2026年7月24日两小时内,对2400次基于真实意图的提示进行4个模型(GPT-5.6 Sol、Gemini 3.6 Flash、Perplexity Sonar Pro、Grok 4.5)、5种语言、5个欧洲市场、每轮4次迭代的测试。所有回应均通过不依赖名册的规则级联编码,剔除同名美国城市误检(精确率96.9%,召回率61.7%),校正提示内聚相关性后有效样本量为407。模型在25.8%的回复中命名个人。类别主导:房地产(35.4%)、汽车经销商(32.9%)显著高于保险(9.1%),卡方检验χ²=159.3,p=5.8e-8。模型差异四倍:Grok达38.0%,Gemini仅9.3%。命名行为受引用类型预测,而非引用数量:提及个人网站多2.6分(95%CI +1.4~+3.9),类别门户多4.3分,企业页面引用率相近(44.1% vs 45.5%)。在9对翻译匹配中,英语提问命名率达36.7%,本地语言仅15.6%(比值比3.14,聚类p=0.074,方向明确)。基于公开领英构建的939人名册仅匹配27,293个姓名提及中的128例(0.47%),其中26人曾被命名,名册覆盖率为0.0%至25.4%。以名册为基础的个体可见性测量仅反映模型行为的一小部分且代表性不足。

原文摘要 · Abstract (English)

Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealerships 32.9% against insurance 9.1% (chi-square 159.3, p = 5.8e-8 after correction). Models differ four-fold, from Grok 38.0% to Gemini 9.3%. Citation type predicts naming and citation volume does not: naming responses cite the individual's own site 2.6 points more often (95% CI +1.4 to +3.9) and category portals 4.3 points more often, and cite firm-owned pages at the same rate (44.1% against 45.5%). On nine matched translation pairs, English prompts named an individual in 36.7% of responses against 15.6% for the same question in the local language (OR 3.14, clustered p = 0.074, so the direction is clear and the design cannot close it). A 939-person roster built from public LinkedIn search matched 128 of 27,293 name-shaped mentions (0.47%), 26 of the 939 people were ever named, and the roster-derived rates of 0.0% to 25.4% measure that overlap. Roster-based measurement of individual AI visibility sees a small and unrepresentative slice of what models do.

AI可见性命名行为多语言模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。