大模型代理对话时会互相模仿,导致任务失败,且普遍存在于主流模型中。
Echoing: Identity Failures when LLM Agents Talk to Each Other

- 代理间对话因缺乏引导信号,易出现角色混淆的‘回声’现象。
- 实验显示最高70%对话出现回声,即使高级推理模型仍有32.8%未改善。
- 结构化响应可将回声率降至9%,适合开发协作型AI系统者参考。
当基于大语言模型(LLM)的代理自主交互时,一类无法从单个代理表现预测的新失败模式浮现:代理-代理对话(AxA)中的行为漂移。与人类-代理交互不同,后者由人类提供锚定和引导,而AxA缺乏此类稳定信号,使这些失败具有独特性。本文研究一种典型失败——回声现象,即代理放弃原定角色,转而镜像对话伙伴,从而破坏其既定目标。通过在66种AxA配置、4个领域(3个交易类,1个咨询类)及2500+次对话(超过25万次LLM推理)中进行实验,发现回声现象在主流LLM供应商中普遍存在,回声率高达70%(视模型与领域而定)。此外,即使在具备较强推理能力的模型中,回声率仍达32.8%,且推理过程并未显著降低该比例。我们分析提示词设计与对话动态,发现回声随对话轮数增加(超过7轮)而加剧,并非仅由实验设计不佳所致。最后,提出一种协议级缓解策略:通过引入结构化响应,将回声率降至9%。
原文摘要 · Abstract (English)
As large language model (LLM) based agents interact autonomously with one another, a new class of failures emerges that cannot be predicted from single agent performance: behavioral drifts in agent-agent conversations (AxA). Unlike human-agent interactions, where humans ground and steer conversations, AxA lacks such stabilizing signals, making these failures unique. We investigate one such failure, echoing, where agents abandon their assigned roles and instead mirror their conversational partners, undermining their intended objectives. Through experiments across $66$ AxA configurations, $4$ domains (3 transactional, 1 advisory), and $2500+$ conversations (over $250000$ LLM inferences), we show that echoing occurs across major LLM providers, with echoing rates as high as $70\%$ depending on the model and domain. Moreover, we find that echoing is persistent even in advanced reasoning models with substantial rates ($32.8\%$) that are not reduced by reasoning efforts. We analyze prompt, conversation dynamics, showing that echoing arises as interaction grows longer ($7+$ agent turns) and is not merely an artifact of sub-optimal experiment design. Finally, we introduce a protocol-level mitigation where targeted use of structured response reduces echoing to $9\%$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。