用日常照片建模用户兴趣,实现跨场景个性化推荐
VisualLens: Personalization through Task-Agnostic Visual History
- 通过多模态大模型分析用户自拍等视觉历史构建通用画像
- 在两个新数据集上提升推荐效果5-10%,超越GPT-4o 2-5%
- 适用于长历史和未见过的物品类别,适合多场景推荐应用
现有推荐系统依赖用户交互日志或文本信号,但物品级历史难以获取且不适用于多模态推荐。我们提出假设:用户日常照片构成的视觉历史可提供丰富、通用的兴趣洞察,可用于有效个性化。为此,我们设计VisualLens框架,利用多模态大语言模型(MLLMs)从视觉历史中提取、过滤并精炼出用户画像以支持个性化推荐。我们构建了两个新基准数据集Google-Review-V和Yelp-V,包含任务无关的视觉历史。实验表明,VisualLens在Hit@3指标上比当前最优物品级多模态推荐方法提升5-10%,优于GPT-4o 2-5%。进一步分析显示,该方法对不同历史长度均具鲁棒性,尤其擅长处理长历史及未见内容类别。
原文摘要 · Abstract (English)
Existing recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals. However, item-based histories are not always accessible, and are not generalizable for multimodal recommendation. We hypothesize that a user's visual history -- comprising images from daily life -- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization. To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history. VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation. We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10% on Hit@3, and outperforms GPT-4o by 2-5%. Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。