arXiv:2608.07069cs.IRcs.CY2026-08

首次对AI餐饮推荐进行全面市场审计,揭示其严重遗漏与偏见。

Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census

论文配图:Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census
图 1 · 摘自论文原文
  • 通过完整枚举两地4776家场所,评估四大AI系统推荐表现
  • 85.6%场所未被推荐,高评分不决定入口,网站和评论量更重要
  • 推荐结果随机性高,跨系统一致性低,关闭门店仍被推荐

AI助手正成为本地发现的主要接口,但对其推荐的场所知之甚少,尤其在餐饮领域,推荐直接影响收入。我们进行了首个基于完整市场普查的AI场所推荐审计:对巴厘岛Canggu和Ubud两地共4,776家咖啡馆、餐厅和酒吧进行全量枚举,并评估四个生产级AI系统(ChatGPT、Claude、Gemini、Perplexity)在七天内对96个角色化查询生成的2,208条响应。因拥有全市场数据,可测量采样审计无法捕捉的结果:85.6%的场所未被任何系统推荐,其中72.6%为拥有50+评价的成熟场所。可见性呈现双门槛结构:进入推荐列表与文档化相关——评论量(OR 1.64)、自有网站(OR 1.92)、价格信息(OR 1.54)、第三方网络提及(OR 1.44),而评分在此阶段无影响(OR 0.89)。一旦被推荐,评分显著预测首位(OR 1.17)。开放POI数据集(Foursquare)的存在在任一门槛均无提升作用。直接编造罕见(0.08%),但系统多次推荐永久关闭场所93次,表明失效主因是过时而非幻觉。跨系统一致性低(前20名杰卡德指数0.33–0.54)。两周测试-重测显示跨周期相似度与同日重跑相当,说明变动源于采样随机性,非时间漂移。论文发布协议、注册表构建方法及衍生数据。

原文摘要 · Abstract (English)

AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.

AI审计推荐系统数据偏见本地搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。