arXiv:2510.04096cs.IRcs.GT2025-10被引 4

用强化学习让AI生成内容抢排名,还能对付对手策略。

RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback

  • 基于排名反馈训练大模型,无需人工标注数据。
  • 在多个测试中显著超越现有方法,提升排名效果。
  • 能适应不同排名系统和对手策略,通用性强。

竞争性搜索是指文档发布者根据查询结果调整内容以提升排名。近期,发布者越来越多地使用大语言模型(LLM)生成和修改内容以增强竞争力。我们提出一种从排名反馈中学习的强化学习框架(RLRF),通过非人类撰写的偏好数据集训练基于LLM的代理。其目标是在考虑对手策略的前提下优化内容以获得更好排名。我们采用不依赖人工数据的方法生成训练数据。实验表明,所提代理在多个场景下均显著优于已有方案;同时具备对未训练过的排名函数(即分布外)的有效性,且能动态适应对手策略。这些结果验证了强化学习在竞争性搜索中的巨大潜力。

原文摘要 · Abstract (English)

Competitive search is a setting where document publishers modify them to improve their ranking in response to a query. Recently, publishers have increasingly leveraged LLMs to generate and modify competitive content. We introduce Reinforcement Learning from Ranker Feedback (RLRF), a framework that trains LLMs using preference datasets derived from ranking competitions. The goal of a publisher (LLM-based) agent is to optimize content for improved ranking while accounting for the strategies of competing agents. We generate the datasets using approaches that do not rely on human-authored data. We show that our proposed agents consistently and substantially outperform previously suggested approaches for LLM-based competitive document modification. We further show that our agents are effective with ranking functions they were not trained for (i.e., out of distribution) and they adapt to strategic opponents. These findings provide support to the significant potential of using reinforcement learning in competitive search.

强化学习信息检索竞争搜索LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。