arXiv:2410.01055cs.CV2024-10

用全景拼接扩展AR检测视野,实时可视化模型表现

ARPOV: Expanding Visualization of Object Detection in AR with Panoramic Mosaic Stitching

  • 通过全景拼接扩展AR视频视域,融合多帧信息
  • 自动过滤无效帧,提升检测结果上下文完整性
  • 交互式分析工具,适合AR与视觉模型调试者使用

随着增强现实(AR)应用日益复杂和普及,智能功能需理解用户行为与环境。现有AR头显视频存在视角局限和镜头抖动问题,难以捕捉用户全视场。传统目标检测可视化仅限单帧,无法呈现时空上下文。本文提出ARPOV,一种面向AR头显视频的交互式可视化分析工具,利用全景拼接扩展环境视图,并自动剔除低质量帧,以增强模型输出的可解释性。该工具由可视化、机器学习与AR专家协同设计,通过5位领域专家访谈验证了其有效性。

原文摘要 · Abstract (English)

As the uses of augmented reality (AR) become more complex and widely available, AR applications will increasingly incorporate intelligent features that require developers to understand the user's behavior and surrounding environment (e.g. an intelligent assistant). Such applications rely on video captured by an AR headset, which often contains disjointed camera movement with a limited field of view that cannot capture the full scope of what the user sees at any given time. Moreover, standard methods of visualizing object detection model outputs are limited to capturing objects within a single frame and timestep, and therefore fail to capture the temporal and spatial context that is often necessary for various domain applications. We propose ARPOV, an interactive visual analytics tool for analyzing object detection model outputs tailored to video captured by an AR headset that maximizes user understanding of model performance. The proposed tool leverages panorama stitching to expand the view of the environment while automatically filtering undesirable frames, and includes interactive features that facilitate object detection model debugging. ARPOV was designed as part of a collaboration between visualization researchers and machine learning and AR experts; we validate our design choices through interviews with 5 domain experts.

AR可视化目标检测全景拼接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。