研究不可靠反馈如何让工具型大模型反而不如不使用工具
Don't Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents

- 控制变量对比可靠、错误和无反馈三种情况下的表现
- 错误反馈下,部分模型表现低于无反馈的基线(如4.7 F1)
- 强调评估时需设置合理无反馈对照组,避免高估工具价值
工具增强型智能体通常在可靠外部反馈下评估性能提升,但这种提升忽略了关键反事实:当反馈不可靠时,智能体是否应完全放弃任务信息?本研究通过受控匹配循环实验,固定智能体流程、提示、动作空间与解码策略,仅改变观测结果——真实、误导或缺失。在问答与事实验证任务中,持续误导性反馈导致价值反转:原本依赖清洁工具的智能体表现反而劣于无反馈基线。在HotpotQA上,Qwen2.5-7B在清洁检索下达到44.8 F1,无反馈为22.3 F1,但在乱序检索下骤降至4.7 F1。该反转现象在更强的清洁检索和局部合理干扰下依然存在,但若后期能获得清洁证据则可缓解。早期轨迹信号可预测多数失败,但简单拒绝错误证据效果有限,仅当暴露的基线可靠时才有效。结果表明,清洁工具带来的增益可能夸大了工具的实际价值,且必须使用匹配的无反馈基线进行评估。
原文摘要 · Abstract (English)
Tool-augmented agents are typically evaluated by their gains under reliable external feedback. Yet these gains leave open a key counterfactual: when feedback is unreliable, would the agent be better off receiving no task evidence? We study this question with a controlled matched-loop comparison that fixes the agent loop, prompt, action space, and decoding, while varying only the returned observation: faithful, misleading, or absent. Across question answering and fact verification, persistent misleading feedback produces a value inversion: agents that benefit from clean tools can perform worse than the matched no-feedback fallback. On HotpotQA, Qwen2.5-7B reaches 44.8 F1 with clean retrieval and 22.3 F1 with no feedback, but drops to 4.7 F1 under shuffled retrieval. The inversion persists under stronger clean retrieval and locally plausible distractors, but weakens when later clean evidence can repair the trajectory. Early trajectory signals predict many failures, yet simple repairs remain fallback-limited: rejecting bad evidence helps only when the exposed fallback is reliable. These results show that clean-tool gains can overstate tool value, and that matched no-feedback fallback controls are necessary for evaluating tool-augmented agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。