arXiv:2604.17969cs.CV2026-04被引 3

构建3D高斯点云场景下的视角依赖视觉搜索基准,测试智能体在自由视角下的主动感知能力。

E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes

  • 设计5自由度视角控制的视觉搜索任务,要求智能体主动调整视角获取信息。
  • 基于99个高保真3D场景和2014个问答任务,评估模型在多视角下的表现差距。
  • 发现当前大模型在动态视角规划上远逊于人类,凸显主动感知短板。

三维环境中的视觉搜索需要具身智能体主动探索并获取任务相关证据。然而,现有视觉搜索与具身AI基准(如EQA)通常依赖静态观察或受限的自身运动,无法显式评估在真实三维环境中自由5-自由度视角控制下出现的细粒度视角依赖现象,例如垂直视角变化导致的可见性改变、容器内内容揭示、仅从特定角度可辨识的对象属性等。为解决此问题,我们提出E3VS-Bench——一个具身三维视觉搜索基准,要求智能体通过5-自由度视角控制收集视角依赖证据以回答问题。该基准包含99个使用3D高斯点云重建的高保真3D场景和2,014个由问题驱动的任务实例。3D高斯点云实现逼真的自由视角渲染,保留了小文字和细微属性等精细视觉细节,这些在基于网格的模拟器中常被弱化,从而支持设计无法单视角解答的问题,必须通过跨视角主动检查来解决。我们评估多个前沿视觉语言模型(VLMs)并对比其与人类的表现。尽管模型具备强大的二维推理能力,但所有模型与人类之间仍存在显著差距,暴露出在全5-自由度视角变化下主动感知与连贯视角规划的局限性。

原文摘要 · Abstract (English)

Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static observations or constrained egocentric motion, and thus do not explicitly evaluate fine-grained viewpoint-dependent phenomena that arise under unrestricted 5-DoF viewpoint control in real-world 3D environments, such as visibility changes caused by vertical viewpoint shifts, revealing contents inside containers, and disambiguating object attributes that are only observable from specific angles. To address this limitation, we introduce {E3VS-Bench}, a benchmark for embodied 3D visual search where agents must control their viewpoints in 5-DoF to gather viewpoint-dependent evidence for question answering. E3VS-Bench consists of 99 high-fidelity 3D scenes reconstructed using 3D Gaussian Splatting and 2,014 question-driven episodes. 3D Gaussian Splatting enables photorealistic free-viewpoint rendering that preserves fine-grained visual details (e.g., small text and subtle attributes) often degraded in mesh-based simulators, thereby allowing the construction of questions that cannot be answered from a single view and instead require active inspection across viewpoints in 5-DoF. We evaluate multiple state-of-the-art VLMs and compare their performance with humans. Despite strong 2D reasoning ability, all models exhibit a substantial gap from humans, highlighting limitations in active perception and coherent viewpoint planning specifically under full 5-DoF viewpoint changes.

具身智能3D感知视觉搜索高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。