arXiv:2601.20316cs.IRcs.LG2026-01

更短的用户历史反而能实现同样推荐质量,还能大幅降低成本。

Less is More: Benchmarking LLM Based Recommendation Agents

  • 用5到10项历史替代50项,推荐效果不变
  • 50用户实验显示质量分稳定在0.17-0.23之间
  • 适合想降本增效的推荐系统开发者

大型语言模型(LLMs)被广泛用于个性化推荐,普遍认为更长的用户购买历史能带来更好预测。我们通过在REGEN数据集上对GPT-4o-mini、DeepSeek-V3、Qwen2.5-72B和Gemini 2.5 Flash四款先进模型进行系统性基准测试,考察了从5到50项商品的历史长度。在50名用户的受试者内设计实验中,结果表明:随着上下文长度增加,推荐质量无显著提升,质量分数始终维持在0.17至0.23之间。这一发现具有重要实践意义:使用5–10项历史可使推理成本降低约88%,而无需牺牲推荐质量。我们还分析了各服务商的延迟模式,揭示模型特异性行为,为部署提供依据。本研究挑战了“更多上下文更好”的固有认知,为低成本、高效的基于LLM的推荐系统提供了可操作指南。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed for personalized product recommendations, with practitioners commonly assuming that longer user purchase histories lead to better predictions. We challenge this assumption through a systematic benchmark of four state of the art LLMs GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flash across context lengths ranging from 5 to 50 items using the REGEN dataset. Surprisingly, our experiments with 50 users in a within subject design reveal no significant quality improvement with increased context length. Quality scores remain flat across all conditions (0.17--0.23). Our findings have significant practical implications: practitioners can reduce inference costs by approximately 88\% by using context (5--10 items) instead of longer histories (50 items), without sacrificing recommendation quality. We also analyze latency patterns across providers and find model specific behaviors that inform deployment decisions. This work challenges the existing ``more context is better'' paradigm and provides actionable guidelines for cost effective LLM based recommendation systems.

推荐系统LLM应用成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。