用大模型实时生成用户兴趣画像,提升推荐精准度与可解释性。
LLM-Based User Personas for Recommendations at Scale

- 基于大模型实时生成自然语言兴趣画像,融合已有兴趣与新主题。
- 在百亿级用户场景下实现低延迟推理,提升观看时长与满意度。
- 适合追求个性化、可解释推荐系统的工业应用与研究者参考。
大型语言模型(LLMs)凭借其世界知识和推理能力,为推荐系统带来巨大潜力。然而,现有方法多依赖结构化ID或离线处理,限制了语义丰富性、实时适应性和用户可理解性。本文提出一种新型框架,可在大规模商业视频推荐平台中实现实时生成基于LLM的用户兴趣画像。该方法在服务阶段直接生成自然语言形式的用户兴趣描述,通过整合已有兴趣与新主题来平衡探索与利用。为应对百亿用户规模下的在线大模型推理计算挑战,设计了高效架构,采用知识蒸馏、异步推理及基于语义聚类的视频表示输入优化。离线评估、用户研究与线上A/B测试均显示显著提升用户价值。本工作弥合了高层语义理解与工业级推荐之间的鸿沟,推动更动态、可解释、令人满意的个性化体验。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer unprecedented potential for enhancing recommendation systems through their world knowledge and reasoning capabilities. However, existing approaches often rely on structured IDs or offline processing, limiting semantic richness, real-time adaptability, and user-facing interpretability. In this paper, we introduce a novel framework that enables real-time generation of LLM-based user interest personas for a large-scale commercial video recommendation platform. Our method generates natural-language user interest personas that address the exploitation-exploration trade-off by combining the summarization of existing interests with novel topics, directly during serving. To overcome the computational challenges of online LLM inference at a billion-user scale, we design a cost-efficient architecture leveraging knowledge distillation, asynchronous inference, and input optimization via semantically clustered video representations. Extensive offline evaluations, user studies, and live A/B tests demonstrate significant improvements in viewer value. This work bridges the gap between high-level semantic understanding and industrial-scale recommendation, paving the way for more dynamic, explainable, and satisfying personalized experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。