arXiv:2603.14669cs.AI2026-03被引 1

让智能体通过渲染视角来推理视野与遮挡关系。

RenderMem: Rendering as Spatial Memory Retrieval

  • 用3D场景渲染代替固定视觉存储,按查询生成对应视角图像。
  • 在AI2-THOR中对视角依赖任务的可见性推理提升显著。
  • 兼容现有视觉语言模型,无需修改架构,适合具身智能研究者。

具身推理本质上依赖视角:可见性、遮挡和可达性取决于智能体的位置。然而,现有空间记忆系统通常仅存储多视角观测或物体中心的抽象信息,难以进行显式的几何推理。我们提出RenderMem,一种将渲染作为3D世界表征与空间推理之间接口的空间记忆框架。不同于存储固定观测,RenderMem维护一个3D场景表示,并根据查询所隐含的视角,渲染出条件化的视觉证据。这使得智能体能从任意视角直接推理视线、可见性和遮挡关系。RenderMem完全兼容现有视觉语言模型,无需修改标准架构。在AI2-THOR环境中的实验表明,其在视角依赖的可见性与遮挡查询任务上持续优于先前的记忆基线。

原文摘要 · Abstract (English)

Embodied reasoning is inherently viewpoint-dependent: what is visible, occluded, or reachable depends critically on where the agent stands. However, existing spatial memory systems for embodied agents typically store either multi-view observations or object-centric abstractions, making it difficult to perform reasoning with explicit geometric grounding. We introduce RenderMem, a spatial memory framework that treats rendering as the interface between 3D world representations and spatial reasoning. Instead of storing fixed observations, RenderMem maintains a 3D scene representation and generates query-conditioned visual evidence by rendering the scene from viewpoints implied by the query. This enables embodied agents to reason directly about line-of-sight, visibility, and occlusion from arbitrary perspectives. RenderMem is fully compatible with existing vision-language models and requires no modification to standard architectures. Experiments in the AI2-THOR environment show consistent improvements on viewpoint-dependent visibility and occlusion queries over prior memory baselines.

空间记忆具身智能3D渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。