发现语言智能体的技能使用存在‘推理暗室’现象,实际依赖与宣称不符。
Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents

- 通过干预技能内容与身份,对比有无技能时的决策差异,测量真实影响。
- 多数模型宣称使用技能但实际依赖变化大,出现沉默采纳或表演性使用。
- 现有检测方法无法识别真正依赖技能的决策,适合审计类研究者参考。
可复用技能正成为扩展语言智能体任务能力的标准接口。然而评估者通常基于可见推理或智能体自述推断技能使用,这些信号反映的是表象而非技能是否真正改变了决策。本文提出是否存在‘推理暗室’——即宣称使用技能与干预测量的实际影响之间的系统性差距。为此,我们构建了BACKTRACE评估框架:为每个带技能条件的回答生成无技能对照,干预技能含义、表述、身份、内容和分配,并在答案提交后才获取归因。在此基础上开发了BACKROOMBench测试平台,覆盖控制逻辑与竞赛数学领域,多种技能条件,单智能体与多智能体场景,以及多样模型家族。评估发现普遍存在溯源失效:模型在不同域中,宣称技能使用稳定,但因果依赖与净效用显著波动,产生静默采纳与表演性使用。行为效应更受程序内容影响,而自述归因则强烈响应外部痕迹的存在。基于直接声明、文本提及、轨迹相似性及大模型判断的观测检测器,均无法准确识别真正依赖技能的决策。在多智能体系统中,技能影响可在通信中延续,即使其来源已消失;而无技能团队仍会虚构从未提供的技能与来源。这些结果确立了‘推理暗室’为普遍的AI溯源问题,其审计必须依赖主动干预。
原文摘要 · Abstract (English)
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's own attribution. These signals show what the agent appears to use, not whether the skill changed its decision. We ask whether skill-augmented agents exhibit a \textbf{Reasoning Backroom}, a systematic gap between stated skill use and intervention-measured influence. We introduce BACKTRACE, an evaluation framework that pairs each skill-conditioned answer with a matched no-skill counterfactual, intervenes on skill meaning, wording, identity, content, and assignment, and elicits attribution only after the answer is committed. We instantiate the framework as BACKROOMBench, a verified testbed spanning controlled logic and competition mathematics, multiple skill conditions, single-agent and multi-agent settings, and diverse model families. Our evaluation reveals a pervasive provenance failure. Across models and domains, stated skill use often remains stable while causal reliance and signed utility vary, producing both silent uptake and performative use. Behavioral effects follow procedural content more reliably than displayed skill identity, whereas stated attributions respond strongly to artifact availability. Observational detectors based on direct skill-use claims, text mentions, trace similarity, and an LLM judge do not identify which decisions actually depend on the skill. In multi-agent systems, skill influence can survive communication even after its source is lost, while no-skill teams still name skills and sources that were never supplied. These findings establish the Reasoning Backroom as a general AI provenance problem whose audit requires intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。