同一提问下,不同买家身份会显著改变AI推荐的品牌,尤其影响中端品牌。
Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit
- 通过10种用户角色+8个问题+3种模型配置的实验设计,测试上下文对推荐的影响。
- 中端品牌推荐变化率达75%,而头部品牌保持约80%一致性。
- Anthropic模型受角色影响更大,因其更依赖训练数据而非检索结果生成。
相同提示词‘最佳CRM软件’在不同用户背景下(如个人创业者、企业高管、英国中小企业主)生成的推荐品牌差异显著。本研究在10种人物角色×8个提示×3种模型配置×10次重复的实验中,共采样2000次运行,其中OpenAI模型覆盖全部8个提示,Anthropic Sonnet-4.6/低覆盖仅4个提示。在用户身份描述前置后,推荐集相似度(Jaccard)相比同角色基线下降0.12至0.20(95%置信区间均排除零值),且该效应呈明显重要性分层:头部品牌推荐一致性达约80%,而中端品牌推荐变化最高达75%。Anthropic模型的效应点估计值高于OpenAI,尽管其置信区间与前者部分重叠;该差异与Anthropic模型更多采用无检索证据生成方式一致(43%-52%推荐无观测检索依据,而OpenAI为8%-29%)。任何对AI品牌认知的测量都必须考虑提问者角色:相同提示因提问者身份不同导致推荐结果实质性差异,跨角色聚合测量会掩盖此关键变异性。该效应集中于中端市场,且在最依赖先验知识的生成路径中最强,表明角色敏感性随模型更依赖训练数据和上下文整合而增强。
原文摘要 · Abstract (English)
The same prompt -- "best CRM software" -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB owner. We audit how strongly that contextual variation reshapes which brands the model recommends. The audit samples 2,000 runs over a design space of 10 personas x 8 prompts x 3 model configurations x N=10 reps, with the two OpenAI cells at full 8-prompt coverage and the Anthropic sonnet-4.6 / low cell at 4-prompt coverage. Prefixing the user message with a persona drops the recommendation-set similarity (Jaccard) by Delta = -0.12 to -0.20 relative to a same-persona baseline (clustered 95% CIs exclude zero on all three measured cells; the sonnet cell's CI rests on only 4 prompt clusters and is correspondingly wider). The effect is sharply prominence-stratified: category leaders are persona-resistant (~80% same-brand consistency across personas), but mid-market brands swap up to 75% of the recommendation set as the persona changes. The Anthropic model shows a larger point-estimate effect than the OpenAI configurations, though clustered CIs overlap for the closer contrast (sonnet vs. OpenAI/high); the asymmetry is consistent with Anthropic's more retrieval-unattributed generation route (43-52% recommendations without observed retrieval-layer evidence, vs OpenAI's 8-29%, documented in Jack 2026). Any measurement of AI brand perception must condition on the buyer persona supplying the query: the same prompt produces materially different recommendation sets depending on who the model thinks is asking, and a measurement protocol that aggregates across personas systematically obscures that variation. The effect concentrates at mid-market and is largest on the most priors-reliant generation route in our audit, consistent with persona responsiveness growing as models lean more on training-data priors and richer context integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。