让大模型学会何时信任自身技能,避免盲目思考和错误调用工具。
Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs

- 用向量分离自信心与外部来源可信度,实现更精准的自我评估。
- 延迟调用判断机制减少无效推理,降低错误调用导致的过度自信。
- 适合构建安全可靠的单智能体系统,尤其关注可信决策场景。
随着大型语言模型(LLMs)融入自主代理与复杂工具生态,传统路由策略日益受到上下文污染和‘过度思考’问题困扰。本文认为瓶颈不在于算法能力或技能多样性,而在于缺乏有纪律的元认知治理。我们提出MESA-S(Metacognitive Skills for Agents, Single-agent)框架,将人类认知控制中的延迟评估、知识警惕性与邻近卸载机制转化为单智能体架构。通过将标量置信度转化为分离自信心(参数确定性)与源可信度(对外部过程的信任)的向量,MESA-S引入延迟过程探测机制与元认知技能卡,使技能可用性认知与其高开销执行解耦。在原生基于Gemini 3.1 Pro运行的上下文静态基准测试中,初步结果表明,显式编程信任溯源与延迟升级可缓解供应链漏洞,剪枝不必要的推理循环,并防止因调用导致的置信度虚高。该架构为可靠、具备知识警惕性的单智能体协调提供了科学审慎且行为锚定的一步。
原文摘要 · Abstract (English)
As large language models (LLMs) transition into autonomous agents integrated with extensive tool ecosystems, traditional routing heuristics increasingly succumb to context pollution and "overthinking". We argue that the bottleneck is not a deficit in algorithmic capability or skill diversity, but the absence of disciplined second-order metacognitive governance. In this paper, our scientific contribution focuses on the computational translation of human cognitive control - specifically, delayed appraisal, epistemic vigilance, and region-of-proximal offloading - into a single-agent architecture. We introduce MESA-S (Metacognitive Skills for Agents, Single-agent), a preliminary framework that shifts scalar confidence estimation into a vector separating self-confidence (parametric certainty) from source-confidence (trust in retrieved external procedures). By formalizing a delayed procedural probe mechanism and introducing Metacognitive Skill Cards, MESA-S decouples the awareness of a skill's utility from its token-intensive execution. Evaluated under an In-Context Static Benchmark Evaluation natively executed via Gemini 3.1 Pro, our early results suggest that explicitly programming trust provenance and delayed escalation mitigates supply-chain vulnerabilities, prunes unnecessary reasoning loops, and prevents offloading-induced confidence inflation. This architecture offers a scientifically cautious, behaviorally anchored step toward reliable, epistemically vigilant single-agent orchestration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。