arXiv:2606.11349cs.AIcs.HC2026-06

让智能体学会在关键时刻主动提问,提升复杂任务中的决策准确性。

Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents

论文配图:Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents
图 1 · 摘自论文原文
  • 将提问作为行动选项之一,与导航竞争,实现动态决策
  • 信息请求有效率从50%提升至74%,显著改善中间判断质量
  • 适用于需要多步推理的复杂分类任务,尤其适合高阶语言模型

在分层推理中,错误常源于中间决策点:智能体在缺乏关键信息时仍选择错误分支。我们提出ACTION-RATING,将提问纳入与导航共享的有序动作空间,使提问与行动直接竞争,帮助请求可被观察于中间状态。智能体自评生成两种信息寻求模式:强制性(无可行分支)和机会性(虽有主导候选但仍有残余不确定性)。在Harmonized Tariff Schedule分类任务(3万节点分类体系,三个基准,9个LLM跨4个模型家族)中,观察到从强制性向机会性提问的转变,信息请求有效性(ISE,即求助后下一步正确导航的比例,非最终任务指标)从50%升至74%。三种诊断对比均无法复现此结构。可分离性测试显示,信息寻求模式(模式分裂、ISE排序)在答案质量下降18.8%准确率时仍保持稳定,支持了智能体何时提问与提问质量之间的经验分离。在受控回答通道下,准确率在10位数字任务上提升+16.2%,这代表了更精准定位所能带来的上限,而非部署预期。

原文摘要 · Abstract (English)

In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that it lacks critical information. Rather than treating clarification as an external uncertainty trigger, we propose ACTION-RATING, a formulation that places it inside the agent's action space on a shared ordinal scale with navigation, so that asking competes directly with acting at every decision point and help-seeking becomes observable at intermediate states. Two structurally distinct information-seeking modes emerge from the agent's own ratings: mandatory (no viable branch) and opportunistic (residual uncertainty despite a leading candidate). On Harmonized Tariff Schedule classification (30,000-node taxonomy, three benchmarks, 9~LLMs across 4 families), we observe a regime shift from mandatory to opportunistic clarification, with Information-Seeking Effectiveness (ISE), a local diagnostic defined as the fraction of help interactions followed by a correct next navigation step (not a final-task metric), rising from 50% to 74%. Three diagnostic contrasts fail to reproduce this structure. A separability test shows that the information-seeking pattern (mode split, ISE ranking) persists when answer quality is degraded (-18.8% accuracy), supporting an empirical separation between where an agent seeks help and the quality of the help it receives. Under the controlled answer channel, accuracy gains reach +16.2% at 10-digit; we read this as an upper bound on what better localization could unlock, not a deployment estimate.

语言模型分层推理主动提问智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。