arXiv:2608.30333cs.IRcs.AI2026-08中稿 · RecSys 2026 Worksh…

LLM生成的推荐理由可作解释组件,但不适宜独立做推荐排序。

Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

论文配图:Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation
图 1 · 摘自论文原文
  • 用行为特征引导LLM生成可读推荐理由
  • LLM评分不如监督模型,但理由质量可独立评估
  • 适合需要可解释性的零售推荐场景

下一篮子复购推荐通常被建模为排序任务:基于用户购买历史,系统对曾购商品进行排序以预测可能再次购买的物品。但在实际应用中,排序准确率只是推荐质量的一部分。客户还可能受益于简洁明了的推荐理由,说明为何此时推荐该商品。大语言模型(LLMs)可通过基于特征、人类可读的理由,利用可解释的行为信号提供此类证据。我们构建了涵盖周期性、频率、最近购买时间、用户行为和商品流行度的复购特征,并在两个公开生鲜数据集和一个私有零售数据集上评估了LLMs的表现。研究问题包括:(1) 现成的LLMs能否作为下一代购物篮评分器,相较于启发式和监督排序器?(2) LLM引用的特征是否携带与结果相关的排序信号?针对后者,我们在跨模型特征掩码协议下,比较了LLM引用特征与模型特异性归因方法的性能下降情况。结果显示,LLM评分无法与监督排序器竞争,表明现成的LLMs不应作为独立的复购推荐系统使用。然而,调整提示词和证据表示方式可在某些场景下提升基于结果的特征掩码效果,即使排序性能未提升;该效应具有数据集依赖性,且并不始终优于归因基线。这些发现表明,LLMs更适合作为经验证的解释组件,而非主要排序器,其理由质量应独立于排序准确率进行评估。

原文摘要 · Abstract (English)

Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential way to surface such evidence through feature-based, human-readable rationales grounded in interpretable behavioral signals. We construct repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to heuristic and supervised rankers, and (2) whether LLM-cited features carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking degradation after masking selected features. Our results show that LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used as standalone repurchase recommenders. However, changes in prompt and evidence representation can improve outcome-grounded feature-masking results in some settings even when ranking performance does not improve; the effect is dataset-dependent and does not consistently match attribution baselines. These findings suggest a practical role for LLMs as validated explanation components rather than primary rankers, with rationale quality evaluated separately from ranking accuracy.

推荐系统可解释性LLM应用零售场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。