arXiv:2603.29798cs.CV2026-03被引 1

用物理仿真验证3D场景功能,发现模型误判问题。

SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes

  • 结合语义推理与几何模拟,分步验证交互可行性。
  • 发现合成室内场景普遍存在无法操作的问题。
  • 适合研究具身智能与视觉语言模型的学者。

具身智能依赖支持有意义活动的交互式3D环境,但评估其功能可用性仍是核心挑战。我们提出SceneTeract框架,基于代理特定约束验证3D场景的功能性。核心是融合高层语义推理与底层几何检查的接地验证引擎。SceneTeract将复杂行为分解为原子动作序列,结合显式物理与几何模拟,逐项验证可达性、间隙和可导航性等条件。我们利用该框架对合成室内环境进行深度评估,发现大量基本交互障碍;同时测试前沿视觉-语言模型(VLMs)对功能性的推理能力,揭示其语义置信度与物理可行性之间存在系统性偏差,即使最强模型亦然。最后,我们将SceneTeract作为奖励机制用于VLM后训练,实现几何约束的规模化知识蒸馏。我们发布SceneTeract验证套件及数据集,以弥合感知与物理现实之间的鸿沟。

原文摘要 · Abstract (English)

Embodied AI depends on interactive 3D environments that support meaningful activities for diverse users, yet assessing their functional affordances remains a core challenge. We introduce SceneTeract, a framework that verifies 3D scene functionality under agent-specific constraints. Our core contribution is a grounded verification engine that couples high-level semantic reasoning with low-level geometric checks. SceneTeract decomposes complex activities into sequences of atomic actions and validates each step against accessibility requirements (e.g., reachability, clearance, and navigability) conditioned on an embodied agent profile, using explicit physical and geometric simulations. We deploy SceneTeract to perform an in-depth evaluation of (i) synthetic indoor environments, uncovering frequent functional failures that prevent basic interactions, and (ii) the ability of frontier Vision-Language Models (VLMs) to reason about and predict functional affordances, revealing systematic mismatches between semantic confidence and physical feasibility even for the strongest current models. Finally, we leverage SceneTeract as a reward engine for VLM post-training, enabling scalable distillation of geometric constraints into reasoning models. We release the SceneTeract verification suite and data to bridge perception and physical reality in embodied 3D scene understanding.

具身智能3D场景视觉语言模型功能验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。