为关键基础设施的LLM智能体动态配置最适工具包,提升效率并降低资源消耗。
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations

- 根据任务需求匹配最小必要工具包,避免过度授权
- 在液冷系统中准确率提升至0.715,耗 token 减少48%
- 适用于对安全与成本敏感的关键系统运维场景
大型语言模型(LLM)代理已广泛应用于关键基础设施(MCI)的运行管理。这些代理通常依赖一个“工具包”来决定其可访问信息、可用工具及可执行动作。现有系统对所有任务统一提供完整工具包,可能造成资源浪费。本文聚焦于最优工具包配置的识别,将其视为任务需求与工具包能力之间的资源匹配问题。通过基于系统数学表征的任务分类,以及对信息量和类型提供能力的排序,构建了任务到工具包的映射关系,来源包括文献挖掘与受控代理执行测量。基于该映射,提出一种“映射引导升级”算法:从任务特异工具包起步,仅在自检失败后才扩展至完整权限。在两项代表性任务中评估:在液冷系统中,准确率从0.652提升至0.715,且达到与Reflexion相当的性能,但仅需48%的token;在电网系统中,完整权限仍为最优,但映射式配置提供了更低开销替代方案。结果表明,工具包配置遵循领域相关的精度-成本帕累托前沿,而非普适最优。
原文摘要 · Abstract (English)
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。