arXiv:2604.00912cs.CVcs.MM2026-04

让投影内容与真实场景分清界限,提升空间增强现实的语义理解能力。

ProCap: Projection-Aware Captioning for Spatial Augmented Reality

  • 通过自动分割分离投影与物理场景,明确区分虚拟与真实内容。
  • 在65个场景、18万+投影上构建首个大规模SAR语义数据集RGBP。
  • 采用双标注评估协议,独立评测对真实场景和投影内容的理解。

空间增强现实(SAR)通过投影仪将数字内容投射到真实场景中,无需头戴设备即可实现沉浸式体验。然而,为支持智能交互(如场景推理或回答用户问题),SAR需能语义区分物理场景与投影内容。标准视觉语言模型(VLMs)难以处理这种虚实混淆,常产生误判。为此,本文提出ProCap框架,通过两阶段流程显式解耦投影内容与物理场景:首先利用自动化分割识别虚拟与物理层;其次采用区域感知检索,避免因投影畸变导致的语义歧义。为支持该方法,我们构建了首个大规模SAR语义基准数据集RGBP,包含65个多样化物理场景及超过18万次投影,附带密集且解耦的标注。此外,设计双标注评估协议,使用任务特定标记分别评估对物理场景与投影内容的描述能力。实验表明,ProCap为未来SAR研究提供了稳健的语义基础。代码、预训练模型及数据集已公开于项目主页:https://ZimoCao.github.io/ProCap/。

原文摘要 · Abstract (English)

Spatial augmented reality (SAR) directly projects digital content onto physical scenes using projectors, creating immersive experience without head-mounted displays. However, for SAR to support intelligent interaction, such as reasoning about the scene or answering user queries, it must semantically distinguish between the physical scene and the projected content. Standard Vision Language Models (VLMs) struggle with this virtual-physical ambiguity, often confusing the two contexts. To address this issue, we introduce ProCap, a novel framework that explicitly decouples projected content from physical scenes. ProCap employs a two-stage pipeline: first it visually isolates virtual and physical layers via automated segmentation; then it uses region-aware retrieval to avoid ambiguous semantic context due to projection distortion. To support this, we present RGBP (RGB + Projections), the first large-scale SAR semantic benchmark dataset, featuring 65 diverse physical scenes and over 180,000 projections with dense, decoupled annotations. Finally, we establish a dual-captioning evaluation protocol using task-specific tokens to assess physical scene and projection descriptions independently. Our experiments show that ProCap provides a robust semantic foundation for future SAR research. The source code, pre-trained models and the RGBP dataset are available on the project page: https://ZimoCao.github.io/ProCap/.

空间增强现实视觉语言模型投影分割数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。