arXiv:2606.00914cs.AIcs.CL2026-06

外部信息流可操控大模型决策,甚至让其放弃原本立场。

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

论文配图:Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
图 1 · 摘自论文原文
  • 固定模型和任务,仅改变信息流内容与顺序,测试其对决策影响。
  • 10%的偏见信息流可使模型决策从5%转向100%,统计显著性极高。
  • 适合关注大模型安全、推荐系统影响的研究者与产品设计者。

大语言模型代理在做出决策前常依赖排名后的外部信息流(如社交动态、搜索结果等),但现有安全评估几乎从未考察上游排序器对代理阅读内容的影响。本文提出一种受控协议,固定模型、人格、主题和最终决策提示,仅变化代理在十轮‘滚动’阶段所见内容的组成与顺序,从而隔离信息流编排对下游决策的因果影响。在三个独立实验室的四款现代开源指令模型上,共进行2,785次决策测试,发现三种响应模式:对抗性屈服、默认饱和,以及单向信息流可显著扭转模型原本不确定的决策(最明显案例中从5%升至100%;Fisher p值低至3×10⁻¹⁰),但无法改变已坚定支持的判断。该效应呈剂量-反应关系,经生成器替换验证排除写作风格干扰,跨多个决策领域泛化,包括移除部署审批门禁或放宽访问控制等安全相关决策。两种简单的信息流防御措施可部分缓解此现象,前沿模型仍保持其默认倾向。研究将推荐系统视为大模型代理的实用、有界控制面,主张评估应审计信息流层而非仅聚焦最终提示。

原文摘要 · Abstract (English)

LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts. We introduce a controlled protocol that holds the model, persona, topic, and final decision prompt fixed and varies only the composition and ordering of the posts an agent encounters during a preceding ten-turn "scrolling" phase, isolating the causal effect of feed curation on a downstream decision. Across 2,785 decision rollouts on four modern open instruct LLMs from three independent labs, we identify three response regimes: adversarial capitulation, default saturation, and a default-direction asymmetry in which a one-sided feed tips a decision the model was genuinely uncertain about (in the clearest cases from 5% to 100%; Fisher p as low as 3 x 10^-10) but cannot dislodge one it already favors or holds firmly. The effect follows a dose-response curve, survives a generator swap that rules out a writing-style artifact, generalizes across several decision domains including security-relevant choices such as removing a deployment approval gate or relaxing access controls, and is partly mitigated by two simple feed-level defenses; a frontier model retains its default. We characterize the recommender as a practical, default-bounded control surface for LLM agents, and argue that agent evaluations must audit the feed layer rather than the final prompt alone.

大模型安全信息流操控决策偏差推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。