建模用户干预模式,让网页代理更懂何时该听、何时该停。
Modeling Distinct Human Interaction in Web Agents
- 基于真实交互数据识别出四种用户参与模式
- 用语言模型预测干预行为,准确率提升超60%
- 实测用户评价代理有用性提高36.8%
尽管自主网页代理发展迅速,人类在任务执行中仍需介入以调整偏好和纠正行为。然而,现有系统缺乏对人类干预时机与原因的系统理解,常在关键节点盲目推进或提出冗余确认。本文提出建模人类干预的任务,构建了包含400条真实用户网页导航轨迹、超过4200次交互动作的CowCorpus数据集。分析发现用户与代理的互动可分为四类:放手监督、主动监督、协作解题、全程接管。基于这些模式,我们训练语言模型预测用户干预意图,在基础模型上实现61.4%-63.4%的准确率提升。进一步将该模型部署于实时网页代理,用户研究表明代理有用性评分提升36.8%。结果表明,结构化建模人类干预可显著提升代理的适应性与协作能力。
原文摘要 · Abstract (English)
Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks unfold. However, current agentic systems lack a principled understanding of when and why humans intervene, often proceeding autonomously past critical decision points or requesting unnecessary confirmation. In this work, we introduce the task of modeling human intervention to support collaborative web task execution. We collect CowCorpus, a dataset of 400 real-user web navigation trajectories containing over 4,200 interleaved human and agent actions. We identify four distinct patterns of user interaction with agents -- hands-off supervision, hands-on oversight, collaborative task-solving, and full user takeover. Leveraging these insights, we train language models (LMs) to anticipate when users are likely to intervene based on their interaction styles, yielding a 61.4-63.4% improvement in intervention prediction accuracy over base LMs. Finally, we deploy these intervention-aware models in live web navigation agents and evaluate them in a user study, finding a 36.8% increase in user-rated agent usefulness. Together, our results show structured modeling of human intervention leads to more adaptive, collaborative agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。