arXiv:2501.07212cs.IR2025-01被引 7

让推荐系统提前规划未来,按目标自动生成个性化内容序列。

Future-Conditioned Recommendations with Multi-Objective Controllable Decision Transformer

  • 基于多目标可控决策变压器,直接指定未来目标生成推荐序列
  • 在多种目标下保持竞争力,无需重复训练即可适应新目标
  • 适合需要长期规划、多目标平衡的推荐场景

实现长期成功是推荐系统的核心目标,要求策略能够预见并塑造决策对未来用户满意度的影响。当前方法面临两大挑战:其一,推荐决策的未来影响难以观测,无法通过即时指标直接优化;其二,不同目标间常存在冲突,如准确率与多样性之间的权衡。现有策略陷入“训练-评估-重训练”的循环,随目标变化而愈发繁琐。为此,我们提出一种未来条件下的多目标可控推荐策略,可直接设定未来目标,使模型自回归生成符合目标的物品序列。我们提出多目标可控决策变压器(MocDT),一种利用大规模离线数据进行离线强化学习的模型,能自主学习从多个目标到物品序列的映射关系。因此,在推理阶段可针对任意指定目标生成推荐。实验表明,该策略能在不同目标下生成相应序列,且性能在各类目标下均与当前先进方法相当。

原文摘要 · Abstract (English)

Securing long-term success is the ultimate aim of recommender systems, demanding strategies capable of foreseeing and shaping the impact of decisions on future user satisfaction. Current recommendation strategies grapple with two significant hurdles. Firstly, the future impacts of recommendation decisions remain obscured, rendering it impractical to evaluate them through direct optimization of immediate metrics. Secondly, conflicts often emerge between multiple objectives, like enhancing accuracy versus exploring diverse recommendations. Existing strategies, trapped in a "training, evaluation, and retraining" loop, grow more labor-intensive as objectives evolve. To address these challenges, we introduce a future-conditioned strategy for multi-objective controllable recommendations, allowing for the direct specification of future objectives and empowering the model to generate item sequences that align with these goals autoregressively. We present the Multi-Objective Controllable Decision Transformer (MocDT), an offline Reinforcement Learning (RL) model capable of autonomously learning the mapping from multiple objectives to item sequences, leveraging extensive offline data. Consequently, it can produce recommendations tailored to any specified objectives during the inference stage. Our empirical findings emphasize the controllable recommendation strategy's ability to produce item sequences according to different objectives while maintaining performance that is competitive with current recommendation strategies across various objectives.

推荐系统多目标优化决策变压器离线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。