arXiv:2510.07635cs.AI2025-10

提出安全探索新推荐项的方法,兼顾系统安全与创新内容曝光。

Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning

  • 基于高置信度离线评估的无模型安全策略学习
  • 实验显示该方法几乎总满足安全要求,但过于保守
  • 通过渐进放松安全约束,实现高效安全部署

在真实推荐系统中,新内容持续不断加入。充分展示新动作对提升长期用户参与度至关重要。尽管近期工作基于离线策略学习(OPL)从历史数据训练策略,但现有方法在面对新动作时可能存在安全隐患。本文目标是构建一个能保证安全的前提下探索新动作的框架。为此,我们首先提出安全离线策略梯度(Safe OPG),一种基于高置信度离线评估的无模型安全OPL方法。首次实验表明,Safe OPG几乎总是满足安全要求,而现有方法常严重违反。但结果也揭示其倾向过度保守,凸显安全与探索间的权衡难题。为克服此问题,我们进一步提出部署高效的安全部署策略学习框架,利用安全余量并仅在多次(非频繁)部署中逐步放松安全正则化。该框架在保障推荐系统安全实施的同时,支持对新动作的有效探索。

原文摘要 · Abstract (English)

In many real recommender systems, novel items are added frequently over time. The importance of sufficiently presenting novel actions has widely been acknowledged for improving long-term user engagement. A recent work builds on Off-Policy Learning (OPL), which trains a policy from only logged data, however, the existing methods can be unsafe in the presence of novel actions. Our goal is to develop a framework to enforce exploration of novel actions with a guarantee for safety. To this end, we first develop Safe Off-Policy Policy Gradient (Safe OPG), which is a model-free safe OPL method based on a high confidence off-policy evaluation. In our first experiment, we observe that Safe OPG almost always satisfies a safety requirement, even when existing methods violate it greatly. However, the result also reveals that Safe OPG tends to be too conservative, suggesting a difficult tradeoff between guaranteeing safety and exploring novel actions. To overcome this tradeoff, we also propose a novel framework called Deployment-Efficient Policy Learning for Safe User Exploration, which leverages safety margin and gradually relaxes safety regularization during multiple (not many) deployments. Our framework thus enables exploration of novel actions while guaranteeing safe implementation of recommender systems.

推荐系统安全探索离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。