arXiv:2411.01114cs.AIcs.CL2024-11被引 14

让大模型更智能地解决复杂工程与逻辑问题,同时大幅降低使用成本。

Infant Agent: A Tool-Integrated, Logic-Driven Agent with Cost-Effective API Usage

  • 整合工具与分层管理,增强大模型持续推理能力
  • 在SWE-bench-lite上准确率从0.33%提升至30%
  • 适合需要高效推理和低成本部署的工程应用

尽管大型语言模型(LLMs)表现出色,但仍存在两大局限:难以自主解决真实世界工程问题,且在复杂逻辑推理上表现不足。为此,我们开发了Infant Agent,集成任务感知功能、操作符、分层管理系统和记忆检索机制。这些组件共同使大模型能够进行长时间推理,高效处理多步骤复杂任务,同时显著降低API调用成本。使用Infant Agent后,GPT-4o在SWE-bench-lite数据集上的准确率从0.33%提升至30%,在AIME-2024数学竞赛中从13.3%提升至37%。

原文摘要 · Abstract (English)

Despite the impressive capabilities of large language models (LLMs), they currently exhibit two primary limitations, \textbf{\uppercase\expandafter{\romannumeral 1}}: They struggle to \textbf{autonomously solve the real world engineering problem}. \textbf{\uppercase\expandafter{\romannumeral 2}}: They remain \textbf{challenged in reasoning through complex logic problems}. To address these challenges, we developed the \textsc{Infant Agent}, integrating task-aware functions, operators, a hierarchical management system, and a memory retrieval mechanism. Together, these components enable large language models to sustain extended reasoning processes and handle complex, multi-step tasks efficiently, all while significantly reducing API costs. Using the \textsc{Infant Agent}, GPT-4o's accuracy on the SWE-bench-lite dataset rises from $\mathbf{0.33\%}$ to $\mathbf{30\%}$, and in the AIME-2024 mathematics competition, it increases GPT-4o's accuracy from $\mathbf{13.3\%}$ to $\mathbf{37\%}$.

大模型推理优化成本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。