arXiv:2508.04102cs.CV2025-08中稿 · MMSys 2026被引 1

用AR打造可复现的视觉评估新平台,让模型表现更贴近真实感知。

AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models

  • 基于增强现实构建模块化评估平台,支持跨平台数据采集与交互任务。
  • 15人用户研究验证其有效性,发现传统指标忽略的感知缺陷。
  • 适合关注模型真实表现、需快速测试视觉效果的研究者使用。

定量评估指标在计算机视觉模型评价中占据核心地位,但常因协议不一致和真实标签噪声而无法反映实际性能。尽管视觉感知研究可补充这些指标,但通常需要端到端系统,实现耗时且难以复现。本文系统梳理了CV模型评估的关键挑战,并提出ARCADE——一个利用增强现实(AR)技术的评估平台,实现便捷、可复现、以人为主导的评估。ARCADE采用模块化架构,支持跨平台数据收集、可插拔模型推理和交互式AR任务,兼顾量化指标与视觉感知评估。通过15名参与者组成的用户研究及两个典型任务(深度估计与光照估计)的案例分析,证明该平台能揭示传统指标忽视的感知瑕疵。同时评估了其易用性与实时性能,证实其作为可靠实时平台的灵活性。

原文摘要 · Abstract (English)

Quantitative metrics are central to evaluating computer vision (CV) models, but they often fail to capture real-world performance due to protocol inconsistencies and ground-truth noise. While visual perception studies can complement these metrics, they often require end-to-end systems that are time-consuming to implement and setups that are difficult to reproduce. We systematically summarize key challenges in evaluating CV models and present the design of ARCADE, an evaluation platform that leverages augmented reality (AR) to enable easy, reproducible, and human-centered CV evaluation. ARCADE uses a modular architecture that provides cross-platform data collection, pluggable model inference, and interactive AR tasks, supporting both metric and visual perception evaluation. We demonstrate ARCADE through a user study with 15 participants and case studies on two representative CV tasks, depth and lighting estimation, showing that ARCADE can reveal perceptual flaws in model quality that are often missed by traditional metrics. We also evaluate ARCADE's usability and performance, showing its flexibility as a reliable real-time platform.

计算机视觉AR评估感知研究模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。