分析大模型推荐应用的逻辑,发现其推荐不透明且不稳定。
Evaluating LLM-Based Mobile App Recommendations: An Empirical Study
- 从大模型输出中提炼出16个通用推荐标准
- 高排名应用结果一致,但越往下差异越大
- 对用户指令反应不一,适合研究对话式推荐
大型语言模型(LLMs)正被广泛用于通过自然语言提示推荐移动应用,提供比关键词搜索更灵活的替代方案。然而,这些推荐背后的推理过程仍不透明,引发对其一致性、可解释性以及与传统应用商店优化(ASO)指标对齐程度的担忧。本文对主流通用型大模型在生成、解释和排序应用推荐时的表现进行了实证分析。主要贡献包括:(i) 从大模型输出中提取出16个可泛化的排序标准;(ii) 构建系统化评估框架,分析推荐的一致性和对显式排序指令的响应能力;(iii) 发布可复现的实验包,支持未来基于AI的推荐系统研究。研究发现,大模型依赖于广泛但零散的排序标准,与标准ASO指标仅部分对齐。虽然高位推荐结果在多次运行中保持一致,但随着排名深度增加和搜索条件细化,变异性上升。大模型对显式指令的敏感度差异显著——从大幅调整到几乎不变的输出不等,揭示了其在对话式应用发现中的复杂推理机制。研究结果旨在帮助终端用户、应用开发者和推荐系统研究人员应对新兴的对话式应用发现环境。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to recommend mobile applications through natural language prompts, offering a flexible alternative to keyword-based app store search. Yet, the reasoning behind these recommendations remains opaque, raising questions about their consistency, explainability, and alignment with traditional App Store Optimization (ASO) metrics. In this paper, we present an empirical analysis of how widely-used general purpose LLMs generate, justify, and rank mobile app recommendations. Our contributions are: (i) a taxonomy of 16 generalizable ranking criteria elicited from LLM outputs; (ii) a systematic evaluation framework to analyse recommendation consistency and responsiveness to explicit ranking instructions; and (iii) a replication package to support reproducibility and future research on AI-based recommendation systems. Our findings reveal that LLMs rely on a broad yet fragmented set of ranking criteria, only partially aligned with standard ASO metrics. While top-ranked apps tend to be consistent across runs, variability increases with ranking depth and search specificity. LLMs exhibit varying sensitivity to explicit ranking instructions - ranging from substantial adaptations to near-identical outputs - highlighting their complex reasoning dynamics in conversational app discovery. Our results aim to support end-users, app developers, and recommender-systems researchers in navigating the emerging landscape of conversational app discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。