用强化学习融合多源文本,生成可解释的用户画像。
High Fidelity Textual User Representation over Heterogeneous Sources via Reinforcement Learning
- 通过点击、申请等行为信号作为奖励,强化学习自动提炼关键信息。
- 在领英多产品上实验,显著提升核心业务指标。
- 无需人工标注,适合实时推荐系统中的大模型应用。
大规模求职平台的个性化推荐需要基于多种异构文本数据建模用户,包括个人资料、专业信息和搜索日志。随着推荐系统广泛采用大语言模型(LLMs),从异构来源生成统一、可解释且简洁的表示变得至关重要,尤其在对延迟敏感的在线环境中。本文提出一种新型强化学习框架,为每位用户合成统一的文本表示。该方法以隐式用户交互信号(如点击、申请)为主要奖励,用于提炼关键信息;同时引入基于规则的奖励,确保输出格式与长度符合要求。在领英(LinkedIn)多个产品上的大量离线实验表明,该方法显著提升了关键下游业务指标。本工作提供了一种实用、无标注、可扩展的解决方案,能够生成直接兼容大模型系统的可解释用户表示。
原文摘要 · Abstract (English)
Effective personalization on large-scale job platforms requires modeling members based on heterogeneous textual sources, including profiles, professional data, and search activity logs. As recommender systems increasingly adopt Large Language Models (LLMs), creating unified, interpretable, and concise representations from heterogeneous sources becomes critical, especially for latency-sensitive online environments. In this work, we propose a novel Reinforcement Learning (RL) framework to synthesize a unified textual representation for each member. Our approach leverages implicit user engagement signals (e.g., clicks, applies) as the primary reward to distill salient information. Additionally, the framework is complemented by rule-based rewards that enforce formatting and length constraints. Extensive offline experiments across multiple LinkedIn products, one of the world's largest job platforms, demonstrate significant improvements in key downstream business metrics. This work provides a practical, labeling-free, and scalable solution for constructing interpretable user representations that are directly compatible with LLM-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。