用大模型生成可指令感知的用户嵌入,提升推荐准确性与鲁棒性。
Instruction-aware User Embedding via Synergistic Language and Representation Modeling
- 融合语言与表征空间,通过轻量适配器处理六类异构数据。
- 对比自回归训练使用户嵌入在多场景下准确率显著提升。
- 适合需要抗噪、跨域泛化的推荐与营销系统使用。
用户表征建模对个性化应用日益重要,但现有方法在跨领域泛化性和噪声行为信号敏感性方面表现不佳。我们提出 InstructUE,一种基于大语言模型(LLMs)的指令感知用户嵌入基础模型,可生成通用且指令感知的用户表征。InstructUE采用多编码器架构,结合轻量适配器,高效处理来自六个不同来源的异构数据,并保留其结构特征。同时,提出一种新颖的对比-自回归训练框架,通过精心构建的 UserQA 数据集连接语言与表征空间。该框架同时利用自回归学习捕捉语言空间中的领域知识,以及对比学习对齐表征空间中的用户-文本嵌入,从而增强用户嵌入的指令感知能力与抗噪性。在真实应用场景的大量实验表明,InstructUE 在用户预测、营销和推荐等多个领域显著优于现有方法。结果表明,指令感知的用户建模能有效实现特定场景下的用户信息指令引导去噪,为更通用、更鲁棒的用户表征学习铺平道路。
原文摘要 · Abstract (English)
User representation modeling has become increasingly crucial for personalized applications, yet existing approaches struggle with generalizability across domains and sensitivity to noisy behavioral signals. We present InstructUE, an instruction-aware user embedding foundation model that leverages large language models (LLMs) to generate general and instruction-aware user representations. InstructUE introduces a multi-encoder architecture with a lightweight adapter that efficiently processes heterogeneous data from six different sources while preserving their structural characteristics. Additionally, it proposes a novel contrastive-autoregressive training framework that bridges language and representation spaces through a curated UserQA dataset. The contrastive-autoregressive training framework simultaneously leverages autoregressive learning to capture domain knowledge in language space and contrastive learning to align user-text embeddings in representation space, thereby enhancing the instruction-awareness and noise-robustness of user embeddings. Through extensive experiments on real-world applications, we demonstrate that InstructUE significantly outperforms existing methods across multiple domains including user prediction, marketing, and recommendation scenarios. Our results show that instruction-aware user modeling can effectively achieve instruction-guided denoising of user information in specific scenarios, paving the way for more generalizable and robust user representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。