让自动化工具会问问题,解决用户指令不清导致的执行失败
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
- 设计可交互追问的GUI导航机制,动态补全缺失信息
- 在模糊任务下,能问问题的代理性能恢复至正常水平
- 适合需要高可靠性的智能设备自动化场景
图形化用户界面(GUI)自动化代理正成为提升智能设备任务完成效率的重要工具。然而,用户在传达任务时常遗漏关键信息,而现有代理无法即时获取补充信息,导致性能下降。为此,我们提出一种支持交互式信息补全的自我修正GUI导航任务。构建了Navi-plus数据集,包含GUI场景下的问答对,并提出双流轨迹评估方法来衡量该能力。实验表明,具备提问能力的代理在面对模糊指令时,可完全恢复其任务执行性能。
原文摘要 · Abstract (English)
Graphical user interfaces (GUI) automation agents are emerging as powerful tools, enabling humans to accomplish increasingly complex tasks on smart devices. However, users often inadvertently omit key information when conveying tasks, which hinders agent performance in the current agent paradigm that does not support immediate user intervention. To address this issue, we introduce a $\textbf{Self-Correction GUI Navigation}$ task that incorporates interactive information completion capabilities within GUI agents. We developed the $\textbf{Navi-plus}$ dataset with GUI follow-up question-answer pairs, alongside a $\textbf{Dual-Stream Trajectory Evaluation}$ method to benchmark this new capability. Our results show that agents equipped with the ability to ask GUI follow-up questions can fully recover their performance when faced with ambiguous user tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。