arXiv:2409.10741cs.SEcs.CL2024-09被引 11

用问答方式引导网页功能探索,无需详细参数即可自动导航。

NaviQAte: Functionality-Guided Web Application Navigation

  • 将网页探索转化为问答任务,用多模态输入理解上下文
  • 在两个数据集上分别达44.23%和38.46%成功率,较现有方法提升15%~33%
  • 适合动态网页测试场景,尤其适用于功能泛化需求

端到端网页测试因需探索多样化的网页功能而具有挑战性。当前先进方法如WebCanvas未设计用于广泛的功能探索,依赖具体详细的任务描述,限制了其在动态网页环境中的适应性。我们提出NaviQAte,将网页应用探索建模为问答任务,生成无须详细参数的功能操作序列。该方法采用三阶段框架,利用GPT-4o等大模型处理复杂决策,结合低成本的GPT-4o mini执行简单任务。NaviQAte聚焦功能导向的网页应用导航,整合文本与图像等多模态输入以增强上下文理解。在Mind2Web-Live与Mind2Web-Live-Abstracted数据集上的评估显示,其用户任务导航成功率达44.23%,功能导航成功率为38.46%,相较WebCanvas分别提升15%和33%。这些结果验证了该方法在推进自动化网页测试方面的有效性。

原文摘要 · Abstract (English)

End-to-end web testing is challenging due to the need to explore diverse web application functionalities. Current state-of-the-art methods, such as WebCanvas, are not designed for broad functionality exploration; they rely on specific, detailed task descriptions, limiting their adaptability in dynamic web environments. We introduce NaviQAte, which frames web application exploration as a question-and-answer task, generating action sequences for functionalities without requiring detailed parameters. Our three-phase approach utilizes advanced large language models like GPT-4o for complex decision-making and cost-effective models, such as GPT-4o mini, for simpler tasks. NaviQAte focuses on functionality-guided web application navigation, integrating multi-modal inputs such as text and images to enhance contextual understanding. Evaluations on the Mind2Web-Live and Mind2Web-Live-Abstracted datasets show that NaviQAte achieves a 44.23% success rate in user task navigation and a 38.46% success rate in functionality navigation, representing a 15% and 33% improvement over WebCanvas. These results underscore the effectiveness of our approach in advancing automated web application testing.

自动化测试网页导航多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。