不推荐内容,直接生成个性化多模态内容
Generate, Not Recommend: Personalized Multimodal Content Generation
- 用大模型直接生成用户定制的图文内容
- 生成图像与用户历史偏好高度匹配,且预判潜在兴趣
- 适合需要创意生成而非筛选的场景
为应对海量网络内容带来的信息过载问题,推荐系统广泛用于为用户筛选和呈现个性化结果。然而,推荐任务本质上局限于现有内容的过滤与选择,缺乏生成新概念的能力,难以完全满足用户需求。本文提出一种新范式:超越内容筛选,直接以多模态形式(如图像)生成个性化内容。为此,我们采用任意对任意的大规模多模态模型(LMMs),通过监督微调与在线强化学习策略训练,使其具备为用户生成定制化下一内容的能力。在两个基准数据集上的实验及用户研究验证了该方法的有效性。值得注意的是,生成的图像不仅与用户历史偏好高度一致,还与潜在未来兴趣相关。
原文摘要 · Abstract (English)
To address the challenge of information overload from massive web contents, recommender systems are widely applied to retrieve and present personalized results for users. However, recommendation tasks are inherently constrained to filtering existing items and lack the ability to generate novel concepts, limiting their capacity to fully satisfy user demands and preferences. In this paper, we propose a new paradigm that goes beyond content filtering and selecting: directly generating personalized items in a multimodal form, such as images, tailored to individual users. To accomplish this, we leverage any-to-any Large Multimodal Models (LMMs) and train them in both supervised fine-tuning and online reinforcement learning strategy to equip them with the ability to yield tailored next items for users. Experiments on two benchmark datasets and user study confirm the efficacy of the proposed method. Notably, the generated images not only align well with users' historical preferences but also exhibit relevance to their potential future interests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。