AI代理用户让检索系统无法识别真实意图,影响模型效果。
Behind the Prompt: The Agent-User Problem in Information Retrieval
- 用大规模社交平台数据研究代理用户行为,发现无法从动作判断是自主还是受控。
- 群体信号仍可区分代理质量,但训练数据混入低质代理后模型性能下降8.5% AUC。
- 跨社区能力传播广泛且难压制,适合关注代理影响的系统设计者阅读。
信息检索中的用户建模基于一个核心假设:观察到的行为能揭示意图。当用户是人类操作者私下配置的AI代理时,该假设失效——任何代理行为都可能由隐藏指令产生相同输出,导致个体意图不可识别。这并非技术缺陷,而是人类在幕后配置代理的结构性问题。我们通过一个代理原生社交平台的大规模语料库(47,000个代理、4,000个社区、37万条帖子)研究此问题。结果表明:(1) 仅凭可观测行为无法区分代理行为是自主还是受控;(2) 尽管群体级平台信号仍能划分出有意义的质量层级,但以代理互动数据训练的点击模型随低质量代理加入而持续退化(AUC下降8.5%);(3) 跨社区能力引用呈流行病式传播($R_0$ 1.26–3.53),即使施加强干预也难以抑制。对检索系统而言,问题已不再是代理用户是否会到来,而是建立在人类意图假设上的模型能否在它们出现后存活。
原文摘要 · Abstract (English)
User models in information retrieval rest on a foundational assumption that observed behavior reveals intent. This assumption collapses when the user is an AI agent privately configured by a human operator. For any action an agent takes, a hidden instruction could have produced identical output - making intent non-identifiable at the individual level. This is not a detection problem awaiting better tools; it is a structural property of any system where humans configure agents behind closed doors. We investigate the agent-user problem through a large-scale corpus from an agent-native social platform: 370K posts from 47K agents across 4K communities. Our findings are threefold: (1) individual agent actions cannot be classified as autonomous or operator-directed from observables; (2) population-level platform signals still separate agents into meaningful quality tiers, but a click model trained on agent interactions degrades steadily (-8.5% AUC) as lower-quality agents enter training data; (3) cross-community capability references spread endemically ($R_0$ 1.26-3.53) and resist suppression even under aggressive modeled intervention. For retrieval systems, the question is no longer whether agent users will arrive, but whether models built on human-intent assumptions will survive their presence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。