arXiv:2601.01857cs.AI2026-01

让大模型代理在真实场景中更稳定,减少工具调用失败

Jenius Agent: Towards Experience-Driven Accuracy Optimization in Real-World Scenarios

  • 通过自适应提示和分层记忆提升长任务执行稳定性
  • 任务完成率最高提升35%,工具调用失败率显著降低
  • 适合需要高可靠性的自动化系统开发者使用

随着大语言模型驱动的智能体系统发展,提升其在上下文理解、工具使用和长周期任务执行方面的能力变得至关重要。然而,现有智能体框架与评测基准对执行层面的行为可见性不足,导致工具调用失败、状态跟踪错误和上下文管理问题难以诊断。本文提出Jenius-Agent,一个基于真实部署经验的系统级智能体框架,集成自适应提示生成、上下文感知的工具编排与分层记忆机制,以增强长周期、工具增强型任务中的执行鲁棒性。此外,我们设计了一种联合评估方法,同时衡量过程保真度、语义正确性和效率。该框架将智能体行为可视化为结构化执行流程,支持对输出指标无法捕捉的失败模式进行系统分析。在Jenius-bench上的实验表明,相较于基线智能体,任务完成率最高提升35%,同时减少令牌消耗、响应延迟及工具调用失败次数。该框架已部署于Jenius(https://www.jenius.cn),提供轻量、可扩展且协议兼容的自主智能体解决方案。

原文摘要 · Abstract (English)

As agent systems powered by large language models (LLMs) advance, improving performance in context understanding, tool usage, and long-horizon execution has become critical. However, existing agent frameworks and benchmarks provide limited visibility into execution-level behavior, making failures in tool invocation, state tracking, and context management difficult to diagnose. This paper presents Jenius-Agent, a system-level agent framework grounded in real-world deployment experience. It integrates adaptive prompt generation, context-aware tool orchestration, and layered memory mechanism to stabilize execution and improve robustness in long-horizon, tool-augmented tasks. Beyond system design, we introduce an evaluation methodology that jointly measures procedural fidelity, semantic correctness, and efficiency. This framework makes agent behavior observable as a structured execution process and enables systematic analysis of failure modes not captured by output-only metrics. Experiments on Jenius-bench show substantial improvements in task completion rate, with up to a 35 percent relative gain over the base agent, along with reduced token consumption, response latency, and tool invocation failures. The framework is already deployed in Jenius ({https://www.jenius.cn}), providing a lightweight and scalable solution for robust, protocol-compatible autonomous agents.

智能体系统大模型应用任务执行优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。