arXiv:2502.11078cs.CL2025-02ACL被引 11

用强化学习动态优化用户画像,持续提升行为预测准确率。

DEEPER Insight into Your User: Directed Persona Refinement for Dynamic Persona Modeling

  • 通过迭代强化学习自动寻找画像优化方向
  • 4轮更新后行为预测误差平均降低32.2%
  • 适合需要持续个性化推荐的场景

为提升推荐系统与用户行为预测等个性化应用,近期研究越来越多地采用大语言模型生成可读的人格画像。在动态现实场景中,有效人格建模需利用流式行为数据持续优化用户画像。然而,现有方法——无论是重新生成画像还是增量扩展新行为——往往难以实现画像质量或未来行为预测准确率的持续提升。为此,我们提出DEEPER,一种新型动态人格建模方法,支持持续画像优化。具体而言,我们通过迭代强化学习框架增强模型的方向搜索能力,使其能自动识别有效更新方向,并利用用户行为与模型预测之间的差异优化画像。在涵盖10个领域共4800名用户的动态画像建模实验中,DEEPER展现出卓越的优化能力,四轮更新后用户行为预测误差平均降低32.2%,优于最佳基线22.92%。

原文摘要 · Abstract (English)

To advance personalized applications such as recommendation systems and user behavior prediction, recent research increasingly adopts large language models (LLMs) for human -readable persona modeling. In dynamic real -world scenarios, effective persona modeling necessitates leveraging streaming behavior data to continually optimize user personas. However, existing methods -whether regenerating personas or incrementally extending them with new behaviors -often fail to achieve sustained improvements in persona quality or future behavior prediction accuracy. To address this, we propose DEEPER, a novel approach for dynamic persona modeling that enables continual persona optimization. Specifically, we enhance the model's direction -search capability through an iterative reinforcement learning framework, allowing it to automatically identify effective update directions and optimize personas using discrepancies between user behaviors and model predictions. Extensive experiments on dynamic persona modeling involving 4800 users across 10 domains highlight the superior persona optimization capabilities of DEEPER, delivering an impressive 32.2% average reduction in user behavior prediction error over four update rounds -outperforming the best baseline by a remarkable 22.92%.

人格建模强化学习个性化推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。