arXiv:2605.00012cs.IRcs.AI2026-05综述被引 1

用强化学习改写搜索片段,可操纵大模型摘要结果。

Exploring LLM biases to manipulate AI search overview

论文配图:Exploring LLM biases to manipulate AI search overview
图 1 · 摘自论文原文
  • 用强化学习训练小模型改写搜索摘要,提升被选中概率。
  • 多数情况下能有效操纵大模型摘要的源选择结果。
  • 揭示摘要选择依赖相对优势,适合研究安全与误导风险者。

现代大型语言模型(LLMs)广泛应用于各类商业场景,尤其在网页搜索系统中生成搜索结果摘要的「大模型摘要系统」中扮演关键角色。这类系统利用大模型从搜索结果中筛选最相关来源并生成答案。已有研究表明,大模型存在多种偏见,而大模型摘要系统在源选择和答案生成阶段均可能受其影响(本文重点关注选择阶段)。本研究旨在探究大模型摘要系统的偏见存在性,并探索利用偏见操纵其输出结果的可能性。我们训练一个小型语言模型,采用强化学习方式重写搜索片段,以增加其被大模型摘要系统青睐的概率。实验设置严格限制策略仅作用于片段内容,避免奖励滥用,反映真实网络搜索环境约束。结果表明,大模型摘要系统确实存在偏见,且强化学习在多数情况下可优化片段内容以操控其输出结果。此外,我们证明大模型摘要的选择基于候选来源间的相对优势而非绝对优势。最后,我们考察了操纵带来的安全风险,发现上下文投毒攻击可能导致不准确或有害结果。

原文摘要 · Abstract (English)

Modern large language models (LLMs) are used in many business applications in general, and specifically in web search systems and applications that generate overviews of search results - LLM Overview systems. Such systems are using an LLM to select most relevant sources from search results and generate an answer to the user's query. It is known from many studies that LLMs have different biases, in LLM Overview application both the source selection and answer generation stages may be affected by the biases of LLMs (here we are focusing mainly on the selection stage). This research is focused on investigating the presence of the biases in LLM Overview systems and on biases exploitation to manipulate LLM Overview results. Here we train a small language model using reinforcement learning to rewrite search snippets to increase their likelihood of being preferred by an LLM Overview. Our experimental setup intentionally restricts the policy to operate only on snippets and limits reward-hacking strategies, reflecting realistic constraints of web search environments. The results prove that LLM Overview systems have biases and that reinforcement learning in most of the cases can optimize snippet's content to manipulate LLM Overview results. We also prove that LLM Overview selections are driven by comparative rather than absolute advantages among candidate sources. In addition, we examine safety aspects of LLM Overview manipulation possibilities and show that context poisoning attacks can lead to inaccurate or harmful results.

大模型偏见搜索摘要强化学习安全攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。