arXiv:2409.10993cs.IR2024-09被引 3

融合图文多模态信息,提升推荐系统交互体验与准确性。

Multi-modal Generative Models in Recommendation System

  • 构建跨模态理解框架,联合分析文本与图像信息
  • 支持用户以图片+文字描述进行精准搜索与个性化推荐
  • 适合需要高沉浸式推荐体验的电商与内容平台

当前推荐系统通常仅接受文本或行为信号作为输入,输出为按相关性排序的产品列表。随着生成式AI的发展,用户期望更丰富的交互方式:例如在视觉搜索中,用户可上传一张图片并用自然语言修改其内容(如‘像图中裙子但颜色为红色’)。同时,用户也希望可视化推荐结果的实际使用场景,如衣物上身效果或家具摆放在房间中的样子。这要求推荐系统能挖掘跨模态(如文本与图像)间的共享与互补信息,并以真实、直观的方式呈现产品。然而,现有系统常独立处理不同模态:文本搜索依赖标题与描述匹配,视觉搜索则基于图像比对。本文主张未来推荐系统应基于多模态理解,充分利用零售商掌握的客户与产品双重信息,实现更优推荐。本章综述了同时利用多种数据模态的推荐系统。

原文摘要 · Abstract (English)

Many recommendation systems limit user inputs to text strings or behavior signals such as clicks and purchases, and system outputs to a list of products sorted by relevance. With the advent of generative AI, users have come to expect richer levels of interactions. In visual search, for example, a user may provide a picture of their desired product along with a natural language modification of the content of the picture (e.g., a dress like the one shown in the picture but in red color). Moreover, users may want to better understand the recommendations they receive by visualizing how the product fits their use case, e.g., with a representation of how a garment might look on them, or how a furniture item might look in their room. Such advanced levels of interaction require recommendation systems that are able to discover both shared and complementary information about the product across modalities, and visualize the product in a realistic and informative way. However, existing systems often treat multiple modalities independently: text search is usually done by comparing the user query to product titles and descriptions, while visual search is typically done by comparing an image provided by the customer to product images. We argue that future recommendation systems will benefit from a multi-modal understanding of the products that leverages the rich information retailers have about both customers and products to come up with the best recommendations. In this chapter we review recommendation systems that use multiple data modalities simultaneously.

多模态推荐系统生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。