构建可调控模糊性的长时任务评估框架,提升智能体自主性。
LHAW: Controllable Underspecification for Long-Horizon Tasks
- 通过删减目标、约束、输入、上下文四维度信息生成可控模糊任务
- 285个任务变体验证了当前智能体在模糊环境下的纠错能力差异
- 首个支持成本敏感评估的长周期任务澄清行为分析框架
能够长期稳定运行的长时工作流智能体是实现真正自主系统的关键。其可靠执行依赖于在模糊情境中主动寻求澄清的能力。然而,现有研究受限于缺乏可扩展、任务无关的框架来系统化构建和评估模糊性影响。为此,我们提出LHAW(长时程增强工作流),一个模块化、数据无关的合成管道,可通过在目标、约束、输入、上下文四个维度以可配置强度移除信息,将任意明确任务转化为可控模糊变体。不同于依赖大模型预测模糊性的方法,LHAW通过实测智能体表现,依据最终状态差异将其分类为结果关键型、发散型或良性。我们发布了基于TheAgentCompany、SWE-Bench Pro和MCP-Atlas的285个任务变体,并进行正式分析,揭示当前智能体在模糊场景中检测、推理与解决不确定性的方式。LHAW提供了首个面向成本敏感的长周期任务澄清行为评估框架,推动可靠自主系统的开发。
原文摘要 · Abstract (English)
Long-horizon workflow agents that operate effectively over extended periods are essential for truly autonomous systems. Their reliable execution critically depends on the ability to reason through ambiguous situations in which clarification seeking is necessary to ensure correct task execution. However, progress is limited by the lack of scalable, task-agnostic frameworks for systematically curating and measuring the impact of ambiguity across custom workflows. We address this gap by introducing LHAW (Long-Horizon Augmented Workflows), a modular, dataset-agnostic synthetic pipeline that transforms any well-specified task into controllable underspecified variants by systematically removing information across four dimensions - Goals, Constraints, Inputs, and Context - at configurable severity levels. Unlike approaches that rely on LLM predictions of ambiguity, LHAW validates variants through empirical agent trials, classifying them as outcome-critical, divergent, or benign based on observed terminal state divergence. We release 285 task variants from TheAgentCompany, SWE-Bench Pro and MCP-Atlas according to our taxonomy alongside formal analysis measuring how current agents detect, reason about, and resolve underspecification across ambiguous settings. LHAW provides the first systematic framework for cost-sensitive evaluation of agent clarification behavior in long-horizon settings, enabling development of reliable autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。