通过挖掘物品隐含意图,提升生成推荐中语义ID的质量。
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation

- 用视觉和文本双路径增强物品表征,融合使用意图
- 大模型提取物品自身隐含动机,不依赖用户偏好
- 仅当生成结果相关时才奖励语义质量,防止虚假流畅
语义ID(SIDs)是生成式推荐中的离散物品词汇,其质量取决于量化前保留的物品证据。在商品推荐中,表面元数据常遗漏潜在使用意图,视觉信息在文本中反映较弱,下游策略学习对生成的SID是否语义有效反馈稀疏。本文提出DeepInterestGR框架,在SID量化前通过两个互补路径增强表征:面向推荐的视觉语言模型(VLM)描述和投影图像嵌入;再利用大模型(LLM)挖掘物品侧的隐含意图——即由产品内容暗示的使用动机,而非个性化用户状态。在策略训练中,引入QARM机制,在标准奖励基础上增加相关性门控的语义质量奖励,仅当生成的SID解码为目标物品时才触发奖励,从而确保语义质量不奖励无关但流畅的预测。在Beauty、Sports、Instruments三个Amazon商品评论类别上的实验表明,DeepInterestGR相比强基线,NDCG@5提升最高达15.1%,NDCG@10提升最高达13.9%。消融实验、分支分析、奖励变体与案例研究支持结论:在量化前融合视觉线索与物品侧意图描述,并结合相关性门控的语义奖励,可有效提升基于SID的生成推荐性能。
原文摘要 · Abstract (English)
Semantic IDs (SIDs) provide the discrete item vocabulary used by generative recommendation, but their quality depends on what item evidence is preserved before quantization. In product recommendation, surface metadata often misses latent usage intent, visual evidence may be only weakly reflected in text, and downstream policy learning provides sparse feedback about whether a generated SID corresponds to a semantically useful item. We introduce \textbf{DeepInterestGR}, an intent-enriched SID framework for generative recommendation. Before SID quantization, \textbf{CMSA} enriches item representations through two complementary evidence paths: recommendation-oriented VLM captions and projected image embeddings. \textbf{DCIM} then uses an LLM to mine item-side intent descriptors -- latent usage motivations implied by product content rather than personalized user states. During policy training over the constructed SIDs, \textbf{QARM} adds a relevance-gated semantic-quality bonus on top of standard SID rewards, applying the bonus only when the generated SID decodes to the target item. Thus, semantic quality cannot reward a fluent but irrelevant item prediction. Experiments on three Amazon Product Review categories (Beauty, Sports, and Instruments) show that DeepInterestGR improves over competitive generative and RL-based baselines, with relative gains of up to \textbf{15.1\%} in NDCG@5 and \textbf{13.9\%} in NDCG@10 over the strongest per-metric baseline. Component ablations, CMSA branch analyses, reward variants, and SID-level case studies support a bounded claim: enriching pre-quantization item evidence with visual cues and item-side intent descriptors, together with relevance-gated semantic rewards, improves SID-based generative recommendation under the evaluated settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。