为机器人时空记忆引入不确定性量化,提升长期任务可靠性。
Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees

- 基于多视角视觉语言模型描述的语义离散度,量化对象级语义不确定性。
- 在固定查询预算下,主动选择高质量视角融合,使不确定性降低37%以上。
- 适用于需要可靠记忆的机器人导航与问答系统,尤其在复杂场景中表现突出。
长时机器人操作依赖时空记忆来记录环境状态并用于后续推理。场景图与检索增强系统将视觉语言模型(VLM)描述与持久3D实体关联,赋予丰富语义。然而,VLM描述存在噪声且视角不一致,现有系统将其视为可信源,缺乏检测不可靠描述的机制。本文提出对象级语义不确定性:通过测量多视角下描述的语义分散程度,识别语义未明确的对象。我们将该不确定性得分集成至先进时空语义记忆系统UQ-DAAAM中,利用其主动选择高质视角,在固定查询预算内融合多视角描述以优化对象表征。同时推导出概率保证:所选高质量候选视角更可能降低不确定性。实验表明,不确定性量化可显著提升具身4D记忆系统的可靠性与有效性。在OC-NaVQA基准上,UQ-DAAAM相比基线实现更大不确定性降低,并取得更优的时空问答性能。
原文摘要 · Abstract (English)
Long-horizon robot operation requires spatio-temporal memory to record the environment state and recall it for downstream reasoning. Scene graphs and retrieval-augmented systems ground VLM descriptions to persistent 3D entities with rich semantic descriptions. However, VLM captions are noisy and viewpoint-inconsistent, and existing systems treat them as an oracle with no mechanism to detect unreliable stored descriptions. We introduce object-level semantic uncertainty for multi-view VLM memory: a score that measures object-centric cross-view semantic scatter of captions and identifies semantically unresolved objects. Then, we include our uncertainty scores in an advanced spatial-semantic memory system, that we dub UQ-DAAAM. UQ-DAAAM uses this score to actively refine uncertain objects under a fixed query budget by selecting high-quality views and fusing the resulting multi-view captions into a single object description. We also derive probabilistic guarantees showing that higher-quality candidate views (as selected by our approach) are more likely to reduce uncertainty. Our experiments show that uncertainty quantification can make embodied 4D memory systems more reliable and more effective. In particular, on the OC-NaVQA benchmark, UQ-DAAAM achieves substantially larger uncertainty reduction and better spatio-temporal question answering performance than baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。