arXiv:2603.19585cs.IR2026-03KDD

用强化学习优化短视频搜索满意度,提升长期用户留存。

SaFRO: Satisfaction-Aware Fusion via Dual-Relative Policy Optimization for Short-Video Search

  • 基于查询行为构建满意度奖励模型,捕捉整体用户体验。
  • 提出双相对策略优化,通过组内与跨批比较高效更新融合策略。
  • 显式建模多任务关系,实现上下文感知的权重自适应。

多任务融合在工业级短视频搜索系统中至关重要,通过整合异构预测信号生成统一排序分数。然而,现有方法主要优化即时互动指标,难以对齐长期用户满意度。虽然强化学习(RL)为满意度优化提供了潜力,但其直接应用于搜索场景面临数据稀疏性和意图约束等挑战。为此,我们提出SaFRO框架,旨在优化短视频搜索中的用户满意度。首先构建一个感知满意度的奖励模型,利用查询级行为代理捕捉超越单个物品交互的整体满意度。随后引入双相对策略优化(DRPO),一种通过组内与跨批次相对偏好比较来高效更新融合策略的方法。此外,设计任务关系感知融合模块,显式建模不同目标间的依赖关系,实现上下文敏感的权重自适应。大规模离线评估和快手平台上的在线A/B测试表明,SaFRO显著优于现有最优基线,在短期排序质量与长期用户留存方面均取得显著提升。

原文摘要 · Abstract (English)

Multi-Task Fusion plays a pivotal role in industrial short-video search systems by aggregating heterogeneous prediction signals into a unified ranking score. However, existing approaches predominantly optimize for immediate engagement metrics, which often fail to align with long-term user satisfaction. While Reinforcement Learning (RL) offers a promising avenue for user satisfaction optimization, its direct application to search scenarios is non-trivial due to the inherent data sparsity and intent constraints compared to recommendation feeds. To this end, we propose SaFRO, a novel framework designed to optimize user satisfaction in short-video search. We first construct a satisfaction-aware reward model that utilizes query-level behavioral proxies to capture holistic user satisfaction beyond item-level interactions. Then we introduce Dual-Relative Policy Optimization (DRPO), an efficient policy learning method that updates the fusion policy through relative preference comparisons within groups and across batches. Furthermore, we design a Task-Relation-Aware Fusion module to explicitly model the interdependencies among different objectives, enabling context-sensitive weight adaptation. Extensive offline evaluations and large-scale online A/B tests on Kuaishou short-video search platform demonstrate that SaFRO significantly outperforms state-of-the-art baselines, delivering substantial gains in both short-term ranking quality and long-term user retention.

短视频搜索强化学习多任务融合用户满意度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。