arXiv:2508.15281cs.IRcs.LG2025-08被引 25

用多模态量化方法生成能适配用户行为的语义编号,提升推荐系统泛化能力。

MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation

  • 分阶段设计多专家混合量化模型,融合跨模态协同与模态特异性信息。
  • 在多个数据集上显著提升新物品和长尾物品的推荐效果,离线指标增益达12.3%。
  • 适合需要处理海量异构内容、动态更新的推荐场景,如电商与短视频平台。

推荐系统传统上使用唯一标识符(ItemID)表示物品,但在大规模、动态变化的物品库和稀疏的长尾数据下,其可扩展性和泛化能力受限。语义ID通过文本、图像等多模态内容映射到共享语义空间,可实现知识迁移,改善对新物品或罕见物品的推荐。然而现有方法面临两大挑战:(1) 如何平衡跨模态协同性与模态特异性;(2) 如何弥合语义表示与真实用户偏好之间的差距。为此,我们提出多模态混合量化(MMQ)框架,包含两阶段训练机制。首先,共享-特异性分词器采用多专家架构,结合模态特异性与共享专家,并使用正交正则化捕捉全面的多模态信息。其次,行为感知微调通过多模态重建损失动态适应下游推荐目标,同时保留模态信息。大量离线实验与在线A/B测试表明,MMQ有效统一了多模态协同性、特异性与行为自适应,为生成式检索与判别式排序任务提供了可扩展且通用的解决方案。

原文摘要 · Abstract (English)

Recommender systems traditionally represent items using unique identifiers (ItemIDs), but this approach struggles with large, dynamic item corpora and sparse long-tail data, limiting scalability and generalization. Semantic IDs, derived from multimodal content such as text and images, offer a promising alternative by mapping items into a shared semantic space, enabling knowledge transfer and improving recommendations for new or rare items. However, existing methods face two key challenges: (1) balancing cross-modal synergy with modality-specific uniqueness, and (2) bridging the semantic-behavioral gap, where semantic representations may misalign with actual user preferences. To address these challenges, we propose Multimodal Mixture-of-Quantization (MMQ), a two-stage framework that trains a novel multimodal tokenizer. First, a shared-specific tokenizer leverages a multi-expert architecture with modality-specific and modality-shared experts, using orthogonal regularization to capture comprehensive multimodal information. Second, behavior-aware fine-tuning dynamically adapts semantic IDs to downstream recommendation objectives while preserving modality information through a multimodal reconstruction loss. Extensive offline experiments and online A/B tests demonstrate that MMQ effectively unifies multimodal synergy, specificity, and behavioral adaptation, providing a scalable and versatile solution for both generative retrieval and discriminative ranking tasks.

推荐系统多模态语义编码行为适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。