用多模态大模型提升推荐系统对图文内容的理解能力
A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems

- 构建三阶段框架,用大模型生成图文描述并转为特征
- 在大规模系统中实现0.35%离线AUC提升,线上指标微增
- 适合想落地多模态大模型的工业级推荐系统团队
传统推荐系统难以充分挖掘多媒体内容中的高维语义信息,制约了用户偏好建模的准确性。尽管多模态大语言模型(MM-LLMs)具备解析复杂数据的强大能力,但其在低延迟、工业级规模架构中的集成仍面临挑战。为此,我们提出一种通用的MM-LLM驱动多媒体理解框架,采用包含内容理解、表征提取与系统集成的三部分架构,基于LLaMA2模型生成描述性标题,并将其作为分词后的类别特征输入系统。实证评估表明,该方法在离线环境下带来0.35%的AUC提升,在线上大规模场景中实现0.02%的指标改善,验证了利用MM-LLMs增强大规模推荐性能的可行性。
原文摘要 · Abstract (English)
Conventional recommendation systems frequently fail to fully exploit the high-dimensional semantic signals inherent in multimedia content, thereby limiting the fidelity of user preference modeling. While Multimodal Large Language Models (MM-LLMs) offer robust mechanisms for interpreting such complex data, their integration into latency-constrained, industrial-scale architectures remains a significant challenge. To address this, we propose a generalized framework for MM-LLM-driven multimedia understanding. Our methodology employs a tripartite architecture encompassing content interpretation, representation extraction, and systematic pipeline integration, instantiated via a LLaMA2-based model that generates descriptive captions subsequently ingested as tokenized categorical features. Empirical evaluation demonstrates the efficacy of this approach, yielding a $0.35\%$ increase in offline AUC and a $0.02\%$ improvement in online metrics at scale, substantiating the practical viability of leveraging MM-LLMs to enhance large-scale recommendation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。