通过三阶段对齐提升生成推荐的多模态理解与用户兴趣挖掘能力。
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation

- 构建视觉-文本-产品编码三重对齐,融合多模态嵌入与深层用户意图
- 在八个生成任务上联合训练,显著减少幻觉并提升泛化性能
- 适合关注多模态推荐、用户兴趣建模与生成式系统优化的研究者
我们提出TriAlignGR,一种统一的多任务多模态生成推荐框架,实现两阶段多模态语义传播:(i) 通过多模态嵌入将视觉语义直接编码至商品标识符(SID);(ii) 通过视觉描述生成任务使模型解码这些语义。现有SID流程存在两个未被充分研究的问题: extbf{SID内容退化(SCD)},即级联编码与残差量化丢弃关键多模态与兴趣层级语义; extbf{SID语义不透明(SSO)},模型自回归生成SID序列却未能真正理解其含义,导致幻觉与泛化差。先前工作仅解决文本-SID对齐,未利用视觉语义与潜在用户兴趣。TriAlignGR通过三个紧密集成组件解决上述问题:(1) extbf{跨模态语义对齐(CMSA)},结合视觉语言模型生成的文本描述与多模态嵌入模型,将图像特征与文本一同编码进SID,确保其自带多模态语义;(2) extbf{多模态深层兴趣挖掘(MDIM)},利用大模型链式思维推理提取潜在用户意图(如“注重效率的生活方式”),超越表面属性,在离散化前丰富SID语义;(3) extbf{三角多任务(TMT)},在单一自回归损失下联合训练八个互补生成任务——包括两项新提出的视觉-语义任务(VisDesc→SID, VisDesc→Title),将VLM生成的图像描述映射到SID与标题,完成SID-文本-图像三角闭环——无需任务专属模块或复杂损失加权。
原文摘要 · Abstract (English)
We introduce TriAlignGR, a unified multitask-multimodal framework for generative recommendation that establishes two-stage multimodal semantic propagation: (i) encoding visual semantics directly into SIDs via multimodal embeddings, and (ii) enabling the model to decode these semantics through visual description tasks. Existing Semantic ID (SID) pipelines suffer from two fundamental but underexplored problems: \textbf{SID Content Degradation (SCD)}, where cascaded encoding and residual quantization discard critical multimodal and interest-level semantics; and \textbf{SID Semantic Opacity (SSO)}, where models autoregressively generate SID sequences without truly comprehending their underlying meaning, leading to hallucination and poor generalization. Prior work addresses at most text-SID alignment, leaving visual semantics and latent user interests entirely unexploited. TriAlignGR resolves both problems through three tightly integrated components: (1)~\textbf{Cross-Modal Semantic Alignment (CMSA)} integrates visual content into SID construction through both VLM-generated textual descriptions and a multimodal embedding model that directly encodes image features alongside text, ensuring that SIDs inherently carry multimodal semantics; (2)~\textbf{Multimodal Deep Interest Mining (MDIM)} leverages LLM Chain-of-Thought reasoning to extract latent user intents (\eg ``productivity-focused lifestyle'' from noise-canceling headphones) beyond surface attributes, enriching SID semantics before discretization; and (3)~\textbf{Triangular Multitask (TMT)} jointly trains on eight complementary generation tasks under a single autoregressive loss -- including two novel visual-semantic tasks (VisDesc$\to$SID, VisDesc$\to$Title) that map VLM-generated image descriptions to SIDs and titles, completing the SID-Text-Image triangle -- without requiring task-specific towers or complex loss weighting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。