用流匹配建模物体在家庭环境中的长期多模态运动规律
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching

- 基于隐空间流匹配,学习物体在3D环境中随时间变化的多模态分布
- 在600多个仿真与真实任务中表现优于现有方法,提升机器人导航效率
- 适合需要理解人类日常行为模式的智能机器人系统研发者
机器人在日常家庭环境中需同时理解三维场景的空间结构与随时间演变的动态特性。人类日常活动使物体位置持续变化,导致机器人难以关联当前观测与历史对象。然而,这些变化具有规律性:由习惯和日常流程形成的时空一致性模式可被机器人学习并利用。为此,我们提出FlowMaps,一种基于隐空间流匹配的模型,用于估计动态物体在未来位置上的多模态分布,支持在连续3D空间中建模物体间依赖关系及其时序演化。该模型能根据过往人类交互预测物体可能的位置变化,并在未见过的相似环境中实现泛化。我们在仿真与真实世界环境中部署了该方法,在超过600个任务中验证其有效性,结果显示,通过连续、多模态的时空分布建模,显著提升了机器人在变化环境中的搜索与导航性能。代码及补充材料见https://fra-tsuna.github.io/flowmaps/。
原文摘要 · Abstract (English)
Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such agents must not only comprehend and navigate spatial layouts, but also reason about how these spaces evolve over time. In particular, humans interact with objects daily, causing them to change position throughout the environment and making it difficult for robots to reliably associate current observations with previously seen objects. However, these interactions are not random: human habits and routines induce spatio-temporally consistent patterns in object locations, which robotic agents can potentially learn and then exploit for downstream tasks such as navigation. To this end, we introduce FlowMaps, a latent flow matching model for estimating multimodal distributions over the future locations of dynamic objects in a continuous 3D space. By learning the implicit dependencies among objects and their temporal evolution, FlowMaps predicts likely changes in object locations conditioned on past human interactions, while supporting generalization across previously unseen environments that share similar object routines. To demonstrate the utility of this method, we deploy FlowMaps in a downstream dynamic Object Navigation task in both simulated and real-world environments. Across more than 600 episodes, FlowMaps outperforms state-of-the-art approaches, showing that modeling object dynamics through continuous, multimodal spatio-temporal distributions improves robotic search and navigation in changing household environments. Code and additional material is available at https://fra-tsuna.github.io/flowmaps/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。