arXiv:2605.04726cs.IR2026-05被引 1

将轻量LLM部署到手机端,实时理解用户意图提升推荐精准度

RecGPT-Mobile: On-Device Large Language Models for User Intent Understanding in Taobao Feed Recommendation

论文配图:RecGPT-Mobile: On-Device Large Language Models for User Intent Understanding in Taobao Feed Recommendation
图 1 · 摘自论文原文
  • 在手机端部署轻量LLM,直接分析用户行为理解意图
  • 相比云端推理,响应速度更快,推荐准确率显著提升
  • 适合需要实时推荐的移动电商场景,推动LLM落地应用

从用户的近期交互行为预测其下一个搜索关键词,是现代电商平台中的关键问题,尤其在用户意图快速变化的场景中。大型语言模型(LLMs)具备强大的语义推理能力,近年来被用于增强下一查询预测的训练数据构建。然而,受限于移动端资源,现有方案多部署在云服务器,导致推理成本高。本文提出RecGPT-Mobile框架,设计了一种轻量级基于LLM的意图理解代理,用于提升移动端电商推荐质量。通过将LLM直接部署在移动设备上,该方法能更快速捕捉用户兴趣变化,并实现实时推荐调整。大量离线分析与在线实验表明,本方法显著提升了推荐结果的准确性,为大规模生产环境中移动端部署LLM提供了可行路径,也为真实世界中下一查询预测系统集成LLM提供了可扩展解决方案。

原文摘要 · Abstract (English)

Predicting a user's next search query from recent interaction behaviors is a critical problem in modern e-commerce systems, particularly in scenarios where user intent evolves rapidly. Large Language Models (LLMs) offer strong semantic reasoning capabilities and have recently been adopted to enhance training data construction for next-query prediction. However, due to resource constraints on mobile devices, existing applications are deployed on cloud servers, resulting in high inference costs. In this paper, we propose RecGPT-Mobile, a framework that designs a lightweight LLM-based intent understanding agent to improve recommendation quality in mobile e-commerce scenarios. By deploying LLMs directly on mobile devices, our approach can capture evolving interests of users more quickly and adjust the recommendation results in real time. Extensive offline analyses and online experiments demonstrate that our method significantly improves the accuracy of recommendation results, laying a practical path for LLM deployment in production-scale recommendation systems on mobile devices, as well as a scalable solution for integrating LLMs into real-world next-query prediction systems.

大模型推荐系统移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。