arXiv:2409.10811cs.SEcs.AI2024-09被引 1

面向虚拟现实应用的零样本可交互界面元素检测框架

Grounded GUI Understanding for Vision-Based Spatial Intelligent Agent: Exemplified by Extended Reality Apps

  • 模仿人类行为先理解场景语义上下文,再进行检测
  • 在反馈引导的验证与反思循环中迭代优化检测结果
  • 无需标注数据即可识别多样异构的可交互元素

近年来,空间计算(即扩展现实,XR)作为一项变革性技术,为用户在多样化虚拟环境中提供了沉浸式交互体验。用户可通过立体三维图形界面(GUI)上的可交互界面元素(IGEs)与XR应用互动。准确识别这些IGEs是自动化测试、有效GUI搜索等软件工程任务的基础。现有针对2D移动应用的IGE检测方法通常依赖大规模人工标注的数据集训练监督型目标检测模型,且仅支持预定义类别(如按钮、滚动条)。然而,由于开放词汇和异构的IGE类别、上下文敏感的可交互性、以及精确的空间感知与视觉-语义对齐需求,此类方法难以适用于XR应用。因此,亟需针对XR应用开展专用的IGE研究。本文提出首个面向虚拟现实应用的零样本上下文感知可交互界面元素检测框架Orienter。Orienter通过模仿人类行为,先观察并理解XR应用场景的语义上下文,再执行检测,并在反馈驱动的验证与反思循环中迭代优化。其包含三个组件:(1) 语义上下文理解,(2) 反思引导的IGE候选检测,(3) 上下文感知的可交互性分类。大量实验表明,Orienter在性能上优于现有最先进的GUI元素检测方法。

原文摘要 · Abstract (English)

In recent years, spatial computing a.k.a. Extended Reality (XR) has emerged as a transformative technology, offering users immersive and interactive experiences across diversified virtual environments. Users can interact with XR apps through interactable GUI elements (IGEs) on the stereoscopic three-dimensional (3D) graphical user interface (GUI). The accurate recognition of these IGEs is instrumental, serving as the foundation of many software engineering tasks, including automated testing and effective GUI search. The most recent IGE detection approaches for 2D mobile apps typically train a supervised object detection model based on a large-scale manually-labeled GUI dataset, usually with a pre-defined set of clickable GUI element categories like buttons and spinners. Such approaches can hardly be applied to IGE detection in XR apps, due to a multitude of challenges including complexities posed by open-vocabulary and heterogeneous IGE categories, intricacies of context-sensitive interactability, and the necessities of precise spatial perception and visual-semantic alignment for accurate IGE detection results. Thus, it is necessary to embark on the IGE research tailored to XR apps. In this paper, we propose the first zero-shot cOntext-sensitive inteRactable GUI ElemeNT dEtection framework for virtual Reality apps, named Orienter. By imitating human behaviors, Orienter observes and understands the semantic contexts of XR app scenes first, before performing the detection. The detection process is iterated within a feedback-directed validation and reflection loop. Specifically, Orienter contains three components, including (1) Semantic context comprehension, (2) Reflection-directed IGE candidate detection, and (3) Context-sensitive interactability classification. Extensive experiments demonstrate that Orienter is more effective than the state-of-the-art GUI element detection approaches.

视觉智能体界面检测扩展现实零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。