实时构建带语义的4D场景图,让机器随时理解任何位置的环境。
Describe Anything Anywhere At Any Moment

- 用批量处理优化局部描述模型,实现高速语义推断。
- 构建分层4D场景图,保持时空一致性且支持实时更新。
- 在多个基准上显著提升问答与任务定位准确率,适合机器人导航等应用。
从增强现实到大规模环境中的机器人自主,计算机视觉与机器人应用需要能同时捕捉几何结构以实现语言对齐并保留语义细节的时空记忆框架。现有方法在生成丰富开放词汇描述时往往牺牲实时性能。为此,我们提出DAAAM(Describe Anything, Anywhere, at Any Moment),一种新型的用于大规模实时4D场景理解的时空记忆框架。DAAAM引入基于优化的前端,通过批量处理从局部描述模型(如Describe Anything Model, DAM)中推断详细语义描述,使在线推理速度提升一个数量级。利用该语义理解构建分层4D场景图(SG),作为全局时空一致的记忆表示。DAAAM在保持实时性的同时构建包含详细几何定位描述的4D SG。我们证明其4D SG可良好对接工具调用代理进行推理与决策。在复杂时空问答任务上,于NaVQA基准评估显示其泛化能力;在SG3D基准上验证序列任务定位能力。我们进一步构建扩展版OC-NaVQA基准用于大规模长时评估。DAAAM在两项任务中均达当前最优:相比最强基线,OC-NaVQA问题准确率提升53.6%,位置误差降低21.9%,时间误差降低21.6%;SG3D任务定位准确率提升27.8%。代码与数据已开源。
原文摘要 · Abstract (English)
Computer vision and robotics applications ranging from augmented reality to robot autonomy in large-scale environments require spatio-temporal memory frameworks that capture both geometric structure for accurate language-grounding as well as semantic detail. Existing methods face a tradeoff, where producing rich open-vocabulary descriptions comes at the expense of real-time performance when these descriptions have to be grounded in 3D. To address these challenges, we propose Describe Anything, Anywhere, at Any Moment (DAAAM), a novel spatio-temporal memory framework for large-scale and real-time 4D scene understanding. DAAAM introduces a novel optimization-based frontend to infer detailed semantic descriptions from localized captioning models, such as the Describe Anything Model (DAM), leveraging batch processing to speed up inference by an order of magnitude for online processing. It leverages such semantic understanding to build a hierarchical 4D scene graph (SG), which acts as an effective globally spatially and temporally consistent memory representation. DAAAM constructs 4D SGs with detailed, geometrically grounded descriptions while maintaining real-time performance. We show that DAAAM's 4D SG interfaces well with a tool-calling agent for inference and reasoning. We thoroughly evaluate DAAAM in the complex task of spatio-temporal question answering on the NaVQA benchmark and show its generalization capabilities for sequential task grounding on the SG3D benchmark. We further curate an extended OC-NaVQA benchmark for large-scale and long-time evaluations. DAAAM achieves state-of-the-art results in both tasks, improving OC-NaVQA question accuracy by 53.6%, position errors by 21.9%, temporal errors by 21.6%, and SG3D task grounding accuracy by 27.8% over the most competitive baselines, respectively. We release our data and code open-source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。