动态场景下,让机器人更智能地选视角、挑记忆,提升问答准确率。
Memory-Guided View Refinement for Dynamic Human-in-the-loop EQA
- 根据视觉重要性动态选择观察视角,避免无效信息积累。
- 在动态与静态场景中均超越现有方法,保持快速推理速度。
- 无需训练,适合真实交互场景中的智能体部署。
具身问答(EQA)传统上在时间稳定的环境中评估,视觉证据可可靠累积。但在动态、有人参与的场景中,人类活动和遮挡带来显著感知非平稳性:任务相关线索短暂且依赖视角,而传统的存-取策略会过度累积冗余证据,增加推理开销。这暴露了两个实际挑战:解决由视角依赖遮挡引起的歧义,以及维持紧凑且实时的证据以实现高效推理。为此,我们提出 DynHiL-EQA,一个包含动态子集(含人类活动与时间变化)和静态子集(时间稳定观测)的人机交互式EQD数据集。针对上述挑战,我们设计了 DIVRR(动态感知视图精炼与相关性引导的自适应记忆选择),一种无需训练的框架,将相关性引导的视图精炼与选择性记忆接纳相结合。通过在存储前验证模糊观测,并仅保留信息量高的证据,DIVRR 在遮挡下提升了鲁棒性,同时保持紧凑内存和快速推理。在 DynHiL-EQA 和 HM-EQA 数据集上的大量实验表明,DIVRR 在动态与静态场景中均持续优于现有基线,且推理效率高。
原文摘要 · Abstract (English)
Embodied Question Answering (EQA) has traditionally been evaluated in temporally stable environments where visual evidence can be accumulated reliably. However, in dynamic, human-populated scenes, human activities and occlusions introduce significant perceptual non-stationarity: task-relevant cues are transient and view-dependent, while a store-then-retrieve strategy over-accumulates redundant evidence and increases inference cost. This setting exposes two practical challenges for EQA agents: resolving ambiguity caused by viewpoint-dependent occlusions, and maintaining compact yet up-to-date evidence for efficient inference. To enable systematic study of this setting, we introduce DynHiL-EQA, a human-in-the-loop EQA dataset with two subsets: a Dynamic subset featuring human activities and temporal changes, and a Static subset with temporally stable observations. To address the above challenges, we present DIVRR (Dynamic-Informed View Refinement and Relevance-guided Adaptive Memory Selection), a training-free framework that couples relevance-guided view refinement with selective memory admission. By verifying ambiguous observations before committing them and retaining only informative evidence, DIVRR improves robustness under occlusions while preserving fast inference with compact memory. Extensive experiments on DynHiL-EQA and the established HM-EQA dataset demonstrate that DIVRR consistently improves over existing baselines in both dynamic and static settings while maintaining high inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。