arXiv:2608.16096cs.IRcs.CL2026-08

评测多跳检索系统时,商业许可和成本被忽视,导致结果不可用。

The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks

  • 审计发现多数领先系统依赖不可商用的嵌入模型,却未披露
  • 2026年前最佳商用模型比基准低2.31个Recall@5点,后被NVIDIA新模型追平
  • 自托管模型可省去高昂索引成本,企业部署需权衡许可与费用

企业在将语言模型接入自有数据时依赖检索技术。现有评测多跳检索系统的基准忽略了两个关键事实:检索底座是否可商用,以及构建成本是多少。在许可方面,领域内主流稠密检索模型NV-Embed-v2采用cc-by-nc-4.0协议,不可商用。我们审计的四个领先MuSiQue系统(HippoRAG-2、PropRAG、SAG、KET-RAG)中,三个依赖该模型获得最佳成绩,但均未说明。性能测试中,我们在同一套MuSiQue基准上对来自八家机构的十三个嵌入模型进行评估,使用置信区间。截至2026年中,存在真实的商业税:最佳商用嵌入模型在Recall@5上落后于基准2.31点(95% CI [0.91, 3.71],p=0.001)。2026年7月16日发布的NVIDIA Nemotron-3-Embed-8B已缩小差距:Recall@5高出0.24点(95% CI [-0.94, +1.43],p=0.69),Recall@10低0.58点(p=0.28),与基准无显著差异。它是唯一同时满足商用许可、可自托管且性能相当的模型;其余满足前两条件的模型仍落后5.2至14.6点。核心结论是付费与免费间的鸿沟:API模型每重索引按令牌计费,自托管则无此开销。在成本方面,五份审计系统中有三份未披露索引成本,而第三方论文中公布的GraphRAG成本相差11倍(2.30美元 vs 24.94美元索引5.64MB语料库一次);扩展至1TB,成本差距达42.8万美元至460万美元。我们的成本模型区分一次性嵌入与持续问答:1TB下嵌入成本为图构建的7.5到900倍,一年1万次查询的问答成本则远低于前者(超350倍)。

原文摘要 · Abstract (English)

Enterprises connect language models to their own data through retrieval. The benchmarks that rank multi-hop retrieval systems leave out two facts a buyer needs before a published number can be used: whether the retrieval backbone may be deployed commercially, and what it costs to build. On licensing: the field's dense-retrieval anchor, NV-Embed-v2, is licensed cc-by-nc-4.0. Of the four leading MuSiQue systems we audit (HippoRAG-2, PropRAG, SAG, KET-RAG), three depend on it for their best numbers and none says so. On performance: we measure thirteen embedders from eight makers on one identical MuSiQue harness with bootstrap confidence intervals throughout. Until mid-2026 there was a real commercial tax: the best commercially-licensed embedder trailed the anchor by 2.31 Recall@5 points (95% CI [0.91, 3.71], p=0.001). NVIDIA's Nemotron-3-Embed-8B, released 2026-07-16, has closed it: +0.24 at Recall@5 (95% CI [-0.94, +1.43], p=0.69), -0.58 at Recall@10 (p=0.28). It matches the anchor, does not beat it, and is the only entrant that is commercially licensed, free to self-host, and indistinguishable from the anchor; every other entrant meeting the first two conditions sits 5.2 to 14.6 points below. The durable finding is the paid-versus-free divide: API embedders charge per token on every re-index, self-hosted ones charge nothing. On cost: three of five audited systems (adding Microsoft's GraphRAG) do not disclose indexing cost, and the only published GraphRAG dollar figures span 11x inside one third-party paper (USD 2.30 vs USD 24.94 to index a 5.64 MB corpus once); extrapolated to 1 TB that undisclosed choice separates roughly USD 428K from $4.6M. Our cost model keeps one-time embedding apart from recurring answering: at 1 TB, embedding sits 7.5x-900x below graph construction, and a year of answering at 10,000 queries/day sits 350x or more below it.

检索系统商业许可成本分析嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。