arXiv:2501.13964cs.CVcs.AI2025-01被引 11

用视觉语言模型评估增强现实场景,发现其识别明显虚拟物效果好,但对融合自然的内容易出错。

Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble

  • 用三款主流VLM模型分析AR场景,基于首个专为评估设计的DiverseAR数据集
  • 识别准确率最高达93%,描述准确率71%,明显虚拟物表现优于无缝融合内容
  • 适合关注AR质量评估、AI理解真实世界能力的研究者和开发者

增强现实(AR)通过叠加虚拟内容提升真实世界体验,但保障其质量、可用性与安全性面临挑战。本研究探索视觉语言模型(VLMs)在自动化评估AR生成场景中的潜力。我们评估了三款前沿商用VLM——GPT、Gemini与Claude,在识别与描述AR场景方面的能力。为此,使用DiverseAR数据集,该数据集是首个专门设计用于评估VLM在多样化复杂度下分析虚拟内容能力的AR数据集。结果显示,VLM普遍能感知并描述AR场景,感知准确率(真阳性率)最高达93%,描述准确率为71%。它们在识别明显虚拟对象(如发光苹果)方面表现良好,但在处理与真实环境无缝融合的内容(如带有逼真阴影的虚拟锅具)时表现不佳。研究揭示了影响VLM性能的关键因素,包括虚拟内容位置、渲染质量及物理合理性。结果表明,VLM在评估AR体验质量方面具有潜力,但也存在局限。

原文摘要 · Abstract (English)

Augmented Reality (AR) enhances the real world by integrating virtual content, yet ensuring the quality, usability, and safety of AR experiences presents significant challenges. Could Vision-Language Models (VLMs) offer a solution for the automated evaluation of AR-generated scenes? Could Vision-Language Models (VLMs) offer a solution for the automated evaluation of AR-generated scenes? In this study, we evaluate the capabilities of three state-of-the-art commercial VLMs -- GPT, Gemini, and Claude -- in identifying and describing AR scenes. For this purpose, we use DiverseAR, the first AR dataset specifically designed to assess VLMs' ability to analyze virtual content across a wide range of AR scene complexities. Our findings demonstrate that VLMs are generally capable of perceiving and describing AR scenes, achieving a True Positive Rate (TPR) of up to 93% for perception and 71% for description. While they excel at identifying obvious virtual objects, such as a glowing apple, they struggle when faced with seamlessly integrated content, such as a virtual pot with realistic shadows. Our results highlight both the strengths and the limitations of VLMs in understanding AR scenarios. We identify key factors affecting VLM performance, including virtual content placement, rendering quality, and physical plausibility. This study underscores the potential of VLMs as tools for evaluating the quality of AR experiences.

AR评估视觉语言模型场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。