arXiv:2508.14912cs.IR2025-08

用多模态自校正对齐提升直播推荐精准度

Multimodal Recommendation via Self-Corrective Preference Alignmen

  • 用大模型生成用户打赏行为的结构化偏好
  • 动态对齐用户偏好与主播多模态特征,提升推荐准确率
  • 适合做直播平台个性化推荐的研究者和工程师

随着直播平台的快速发展,个性化推荐系统在提升用户体验和推动平台收益方面日益关键。直播内容具有动态性和多模态特性(如视觉、音频、文本数据),需联合建模用户行为与多模态特征以捕捉主播属性的演变。然而,依赖单一模态或把多模态作为补充的传统方法难以对齐用户动态偏好与主播多模态属性,限制了推荐的准确性和可解释性。为此,我们提出MSPA(Multimodal Self-Corrective Preference Alignment)框架,包含两个组件:(1) 多模态偏好组合器,利用多模态大语言模型(MLLMs)从用户打赏历史中生成结构化偏好文本和嵌入;(2) 自校正偏好对齐推荐器,将用户偏好与主播多模态特征对齐,提升推荐精度与可解释性。大量实验与可视化表明,MSPA在动态直播场景下显著提升准确率、召回率和文本质量,优于多个基线方法。

原文摘要 · Abstract (English)

With the rapid growth of live streaming platforms, personalized recommendation systems have become pivotal in improving user experience and driving platform revenue. The dynamic and multimodal nature of live streaming content (e.g., visual, audio, textual data) requires joint modeling of user behavior and multimodal features to capture evolving author characteristics. However, traditional methods relying on single-modal features or treating multimodal ones as supplementary struggle to align users' dynamic preferences with authors' multimodal attributes, limiting accuracy and interpretability. To address this, we propose MSPA (Multimodal Self-Corrective Preference Alignment), a personalized author recommendation framework with two components: (1) a Multimodal Preference Composer that uses MLLMs to generate structured preference text and embeddings from users' tipping history; and (2) a Self-Corrective Preference Alignment Recommender that aligns these preferences with authors' multimodal features to improve accuracy and interpretability. Extensive experiments and visualizations show that MSPA significantly improves accuracy, recall, and text quality, outperforming baselines in dynamic live streaming scenarios.

推荐系统多模态直播推荐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。