用三重对齐提升大模型推荐效果,实测性能显著超越当前最佳
Align$^3$GR: Unified Multi-Level Alignment for LLM-based Generative Recommendation
- 通过融合语义与协同信号实现双粒度分词
- 召回率提升17.8%,NDCG提升20.2%,线上测试表现优异
- 适合工业级推荐系统优化与大模型应用落地场景
大型语言模型在利用结构化世界知识和多步推理方面具有显著优势,但将其转化为真实推荐系统时面临语义与行为不匹配的根本挑战。为此,我们提出Align$^3$GR框架,统一了令牌级、行为建模级和偏好级对齐。该方法引入:融合用户-项目语义与协同信号的双粒度分词;双向语义对齐增强行为建模;结合自对弈(SP-DPO)与真实反馈(RF-DPO)的渐进式直接偏好优化策略,实现动态偏好适应。实验表明,该框架在公开数据集上相比最先进基线,召回率@10提升17.8%,NDCG@10提升20.2%;在线A/B测试及工业大规模推荐平台部署中亦取得显著成效。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate significant advantages in leveraging structured world knowledge and multi-step reasoning capabilities. However, fundamental challenges arise when transforming LLMs into real-world recommender systems due to semantic and behavioral misalignment. To bridge this gap, we propose Align$^3$GR, a novel framework that unifies token-level, behavior modeling-level, and preference-level alignment. Our approach introduces: Dual tokenization fusing user-item semantic and collaborative signals. Enhanced behavior modeling with bidirectional semantic alignment. Progressive DPO strategy combining self-play (SP-DPO) and real-world feedback (RF-DPO) for dynamic preference adaptation. Experiments show Align$^3$GR outperforms the SOTA baseline by +17.8% in Recall@10 and +20.2% in NDCG@10 on the public dataset, with significant gains in online A/B tests and full-scale deployment on an industrial large-scale recommendation platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。