通过分层搜索优化企业环境中的智能体框架,显著提升任务完成率。
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

- 按失败模式分层任务,分离可观察与不可观察的搜索环节。
- 在3个企业基准上性能提升20-35个百分点,且跨模型迁移有效。
- 适合需要降低模型与环境不匹配的企业自动化场景。
我们提出StarHarness,一个在固定模型权重前提下演化适配特定环境的智能体框架。该框架可优化提示词、任务表述、工具接口、技能、基于MCP的提供者、子代理结构及代理循环配置。StarHarness通过按基础失败行为分层任务,构建紧凑的演化池,将可观察的搜索任务与隐藏的选择任务分离,并保留未参与演化的任务用于评估泛化能力。在ITBench SRE、EnterpriseOps-Gym ITSM和AutomationBench Finance三个基准上,每个环境经过4-12次采纳变更后,整体性能相比默认框架提升20-35个百分点。这些增益在未参与演化的任务上依然存在,且在GPT与Qwen模型族间无需重新演化即可迁移。轨迹分析表明,性能提升源于接口修复、环境惯例对齐及操作知识压缩,减少了误诊次数并缩短了执行路径。因此,StarHarness为解决工具丰富的企业任务中长期存在的模型-环境不匹配问题提供了实用方案。
原文摘要 · Abstract (English)
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。