arXiv:2601.14691cs.AIcs.CL2026-01被引 3

伪造推理过程可骗过大模型评分,暴露评测体系重大漏洞

Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

  • 通过改写代理的思维链,不改变行为和观察,即可操纵评分
  • 内容型伪造使顶尖视觉语言模型评分误判率最高提升90%(800条轨迹)
  • 现有防御方法仍无法根治漏洞,需基于可观测证据验证推理

大型语言模型(LLMs)正被广泛用作评估代理性能的裁判,尤其在不可验证场景中依赖代理轨迹中的思维链(CoT)推理。该范式隐含假设:代理的思维链真实反映其内部推理与环境状态。我们揭示此假设极脆弱:LLM裁判极易被代理推理痕迹的操控所影响。通过系统性地重写代理的思维链(保持动作和观测不变),我们发现仅改变推理内容即可使当前最先进的视觉语言模型裁判的误报率在800条跨多样化网络任务轨迹上最高上升90%。我们研究了基于风格的呈现改写与基于内容的任务进展信号伪造两种策略,发现内容型操纵始终更有效。评估了提示工程与扩大裁判计算资源等缓解手段,虽能降低但无法完全消除脆弱性。研究结果揭示了基于LLM评估的根本性缺陷,强调需要能将推理主张与可观测证据对齐的评判机制。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories including chain-of-thought (CoT) reasoning. This paradigm implicitly assumes that the agent's CoT faithfully reflects both its internal reasoning and the underlying environment state. We show this assumption is brittle: LLM judges are highly susceptible to manipulation of agent reasoning traces. By systematically rewriting agent CoTs while holding actions and observations fixed, we demonstrate that manipulated reasoning alone can inflate false positive rates of state-of-the-art VLM judges by up to 90% across 800 trajectories spanning diverse web tasks. We study manipulation strategies spanning style-based approaches that alter only the presentation of reasoning and content-based approaches that fabricate signals of task progress, and find that content-based manipulations are consistently more effective. We evaluate prompting-based techniques and scaling judge-time compute, which reduce but do not fully eliminate susceptibility to manipulation. Our findings reveal a fundamental vulnerability in LLM-based evaluation and highlight the need for judging mechanisms that verify reasoning claims against observable evidence.

大模型评测思维链对抗攻击可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。