arXiv:2605.07937cs.CL2026-05被引 2

研究智能体在长任务中何时提问澄清最有效,发现不同信息缺失的提问时机影响巨大。

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

论文配图:Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
图 1 · 摘自论文原文
  • 通过控制实验设计,系统测试了不同信息缺失时提问时机对性能的影响。
  • 目标信息缺失后10%执行阶段就基本无效,输入信息可延迟到50%仍有效。
  • 当前模型普遍不按最优时机提问,适合优化决策策略的研究者参考。

长周期智能体执行包含数百步动作的复杂任务,早期错误假设会引发不可逆错误。当指令不完整时,智能体需决定是否提问及何时提问,但此前缺乏对澄清价值随执行进程变化的度量。本文引入强制注入框架,在四个信息维度(目标、输入、约束、上下文)、三个智能体基准和四种前沿模型(每基准三个;单基准一个;共84个任务变体;6000+次运行)中,精确控制澄清时机。结果表明,与“越早越好”的直觉相反,澄清价值高度依赖缺失信息类型:目标澄清在执行10%后价值近乎消失(pass@3从0.78降至基线),而输入澄清可维持至约50%。任何类型澄清若推迟至中段之后,性能反而低于从不提问。跨模型肯德尔相关系数(0.78-0.87,相同任务覆盖模型间;0.34-0.67,全四模型面板)证实这些时机特征具有强任务内生性。对300次非脚本化会话的补充分析显示,当前所有前沿模型均未在实证最优窗口内提问,策略范围从过度提问(52%会话)到从不提问。这些经验需求曲线为现有理论框架提供了长期缺失的量化基础,并确立了面向时机感知澄清策略的设计目标。代码与数据将公开发布。

原文摘要 · Abstract (English)

Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreversible errors. When instructions are incomplete, the agent must decide not only whether to ask for clarification but when, and no prior work measures how clarification value changes over the course of execution. We introduce a forced-injection framework that provides ground-truth clarifications at controlled points in the agent's trajectory across four information dimensions (goal, input, constraint, context), three agent benchmarks, and four frontier models (three per benchmark; one on a single benchmark only; 84 task variants; 6,000+ runs). Counter to the common intuition that "earlier is always better," we find that the value of clarification depends sharply on what information is missing: goal clarification loses nearly all value after 10% of execution (pass@3 drops from 0.78 to baseline), while input clarification retains value through roughly 50%. Deferring any clarification type past mid-trajectory degrades performance below never asking at all. Cross-model Kendall tau correlations (0.78-0.87 among models sharing identical task coverage; 0.34-0.67 across the full 4-model panel) confirm these timing profiles are substantially task-intrinsic. A complementary study of 300 unscripted sessions reveals that no current frontier model asks within the empirically optimal window, with strategies ranging from over-asking (52% of sessions) to never asking at all. These empirical demand curves provide the quantitative foundation that existing theoretical frameworks require but have lacked, and establish concrete design targets for timing-aware clarification policies. Code and data will be publicly released.

智能体任务规划澄清机制时机优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。