arXiv:2605.09338cs.IR2026-05中稿 · SIGIR 2026 short

用多模态大模型提升推荐系统对图文内容的理解能力

A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems

论文配图:A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems
图 1 · 摘自论文原文
  • 构建三阶段框架,用大模型生成图文描述并转为特征
  • 在大规模系统中实现0.35%离线AUC提升,线上指标微增
  • 适合想落地多模态大模型的工业级推荐系统团队

传统推荐系统难以充分挖掘多媒体内容中的高维语义信息,制约了用户偏好建模的准确性。尽管多模态大语言模型(MM-LLMs)具备解析复杂数据的强大能力,但其在低延迟、工业级规模架构中的集成仍面临挑战。为此,我们提出一种通用的MM-LLM驱动多媒体理解框架,采用包含内容理解、表征提取与系统集成的三部分架构,基于LLaMA2模型生成描述性标题,并将其作为分词后的类别特征输入系统。实证评估表明,该方法在离线环境下带来0.35%的AUC提升,在线上大规模场景中实现0.02%的指标改善,验证了利用MM-LLMs增强大规模推荐性能的可行性。

原文摘要 · Abstract (English)

Conventional recommendation systems frequently fail to fully exploit the high-dimensional semantic signals inherent in multimedia content, thereby limiting the fidelity of user preference modeling. While Multimodal Large Language Models (MM-LLMs) offer robust mechanisms for interpreting such complex data, their integration into latency-constrained, industrial-scale architectures remains a significant challenge. To address this, we propose a generalized framework for MM-LLM-driven multimedia understanding. Our methodology employs a tripartite architecture encompassing content interpretation, representation extraction, and systematic pipeline integration, instantiated via a LLaMA2-based model that generates descriptive captions subsequently ingested as tokenized categorical features. Empirical evaluation demonstrates the efficacy of this approach, yielding a $0.35\%$ increase in offline AUC and a $0.02\%$ improvement in online metrics at scale, substantiating the practical viability of leveraging MM-LLMs to enhance large-scale recommendation performance.

多模态推荐系统大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。