提出在线人物交互生成与感知新任务,解决真实场景下信息滞后问题。
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
- 基于Mamba架构设计带记忆机制的在线框架,实时处理流式数据。
- 在Core4D、OAKINK2等数据集上实现线上生成任务领先性能。
- 适用于机器人、AR/VR等需实时理解交互的场景。
人-物交互(HOI)的感知与生成对机器人、增强现实/虚拟现实(AR/VR)及人类行为理解至关重要。然而,现有方法多在离线环境下建模,可利用完整交互序列信息;而真实场景中,每个时间步仅能获取当前时刻和历史信息,即在线设置。我们发现,离线方法在在线情境下表现不佳。为此,提出两个新任务:在线HOI生成与感知。为此,构建OnlineHOI框架,基于Mamba架构并引入记忆机制,利用Mamba对流数据的强大建模能力及记忆机制高效融合历史信息,在Core4D和OAKINK2的在线生成任务,以及HOI4D的在线感知任务上均取得当前最优结果。
原文摘要 · Abstract (English)
The perception and generation of Human-Object Interaction (HOI) are crucial for fields such as robotics, AR/VR, and human behavior understanding. However, current approaches model this task in an offline setting, where information at each time step can be drawn from the entire interaction sequence. In contrast, in real-world scenarios, the information available at each time step comes only from the current moment and historical data, i.e., an online setting. We find that offline methods perform poorly in an online context. Based on this observation, we propose two new tasks: Online HOI Generation and Perception. To address this task, we introduce the OnlineHOI framework, a network architecture based on the Mamba framework that employs a memory mechanism. By leveraging Mamba's powerful modeling capabilities for streaming data and the Memory mechanism's efficient integration of historical information, we achieve state-of-the-art results on the Core4D and OAKINK2 online generation tasks, as well as the online HOI4D perception task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。