重复提问时,不查网页的模型仍能推荐新品牌,查网页的则很快饱和。
Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources

- 不依赖外部检索的模型在15轮后仍能持续推荐新品牌
- 未检索模型平均保留15-31个品牌,检索模型仅8个且多数已饱和
- 适合研究大模型推荐多样性与检索机制影响的学者
在300个问答单元中(50个问题,6种引擎,每种15次运行),不使用网络搜索的五种引擎在第15轮仍有86%-92%的单元持续推荐未见过的品牌,中位品牌数为15-31个;而唯一启用检索的引擎在第15轮已基本闭合(中位8个品牌,64%仍在增加),与此前四组深度测试中第10轮即饱和的结果一致。在所有测试时间点上,引用领域数量持续增长:四组深度测试在第24轮仍持续扩展,达到Chao2下限估计值的59%-84%;检索引擎44%的广度单元在第15轮仍在添加领域。单次运行可覆盖5轮品牌集的62%-77%;跨引擎统计显示,平均每题生成38个组织,其中中位15个仅出现在单一引擎中。采用精确稀疏化与Chao2丰富度估计,固定名单提取法产生平台曲线,说明名单约束会人为制造饱和假象,开放提取可消除此偏差。
原文摘要 · Abstract (English)
Whether repeated identical buying questions exhaust a language model's brand recommendations depends on retrieval. Across 300 question-engine cells (50 questions, six engines, 15 runs each, open extraction over 1,470 adjudicated organizations), the five engines answering without web search were still adding never-seen brands at run 15 in 86-92% of cells, with median repertoires of 15-31 organizations; the one retrieval-enabled engine closed its list (median 8 organizations, 64% of cells still adding), matching four earlier deep cells where web-search runs saturated by run ten. Cited-domain accumulation keeps rising at every horizon tested: four deep cells were still adding domains at run 24 with 59-84% of the Chao2 lower-bound estimate observed, and 44% of the retrieval engine's breadth cells were still adding domains at run 15. A single run shows 62-77% of the five-run brand set, and across engines the median question draws 38 organizations, of which a median of 15 appear in exactly one engine. Estimators are exact rarefaction and Chao2 richness; a parallel fixed-roster extraction reproduces flat curves on identical responses, so roster-bounded tracking manufactures plateaus that open extraction removes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。