研究发现,智能体工作流易受欺骗性评判影响,哪怕最强模型也可能被误导。
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
- 构建双维度评判框架,区分意图与知识来源
- 强模型在单轮误导反馈后会错误切换正确答案
- 适合关注智能体系统安全与鲁棒性的研究者
智能体工作流依赖多个大语言模型通过反馈机制协作解决问题。然而,评判模型若存在幻觉、偏见或恶意行为,将引入严重漏洞。本文提出一个二维评判行为分析框架,涵盖意图(从建设性到恶意)和知识来源(仅参数化至检索增强)。基于此框架,构建多种评判行为,并开发WAFER-QA基准,使用网络检索证据生成事实支持的对抗性批评,评估工作流鲁棒性。结果表明,即使最强模型在单轮误导性反馈下也会改变正确答案。进一步分析多轮交互中模型预测演变,揭示推理型与非推理型模型的不同行为模式。研究揭示了基于反馈的工作流的根本脆弱性,并为构建更稳健的智能体系统提供指导。
原文摘要 · Abstract (English)
Agentic workflows -- where multiple large language model (LLM) instances interact to solve tasks -- are increasingly built on feedback mechanisms, where one model evaluates and critiques another. Despite the promise of feedback-driven improvement, the stability of agentic workflows rests on the reliability of the judge. However, judges may hallucinate information, exhibit bias, or act adversarially -- introducing critical vulnerabilities into the workflow. In this work, we present a systematic analysis of agentic workflows under deceptive or misleading feedback. We introduce a two-dimensional framework for analyzing judge behavior, along axes of intent (from constructive to malicious) and knowledge (from parametric-only to retrieval-augmented systems). Using this taxonomy, we construct a suite of judge behaviors and develop WAFER-QA, a new benchmark with critiques grounded in retrieved web evidence to evaluate robustness of agentic workflows against factually supported adversarial feedback. We reveal that even strongest agents are vulnerable to persuasive yet flawed critiques -- often switching correct answers after a single round of misleading feedback. Taking a step further, we study how model predictions evolve over multiple rounds of interaction, revealing distinct behavioral patterns between reasoning and non-reasoning models. Our findings highlight fundamental vulnerabilities in feedback-based workflows and offer guidance for building more robust agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。