首个支持开放词汇的4D人体解析框架,推理速度提升93.3%。
OpenHuman4D: Open-Vocabulary 4D Human Parsing
- 用掩码视频目标追踪建立时空对应,跳过全帧分割
- 新掩码验证模块提升新目标识别与追踪鲁棒性
- 4D掩码融合模块实现高效嵌入整合,适合动态场景应用
动态三维人体表征理解在虚拟与扩展现实应用中日益重要。然而,现有方法受限于封闭集数据集和长推理时间,严重制约实用性。本文提出首个同时解决这两问题的4D人体解析框架,实现开放词汇能力并大幅加速推理。基于先进开放词汇3D人体解析技术,本方法将支持扩展至4D人体中心视频,包含三项关键创新:1)采用基于掩码的视频目标追踪,高效建立时空对应关系,无需对所有帧进行分割;2)设计新型掩码验证模块,用于处理新目标识别与缓解追踪失败;3)提出4D掩码融合模块,结合记忆条件注意力与逻辑值均衡,实现鲁棒嵌入融合。大量实验表明,该方法在4D人体中心解析任务上表现优异,相比前序最先进方法提速达93.3%,且不再局限于固定类别。
原文摘要 · Abstract (English)
Understanding dynamic 3D human representation has become increasingly critical in virtual and extended reality applications. However, existing human part segmentation methods are constrained by reliance on closed-set datasets and prolonged inference times, which significantly restrict their applicability. In this paper, we introduce the first 4D human parsing framework that simultaneously addresses these challenges by reducing the inference time and introducing open-vocabulary capabilities. Building upon state-of-the-art open-vocabulary 3D human parsing techniques, our approach extends the support to 4D human-centric video with three key innovations: 1) We adopt mask-based video object tracking to efficiently establish spatial and temporal correspondences, avoiding the necessity of segmenting all frames. 2) A novel Mask Validation module is designed to manage new target identification and mitigate tracking failures. 3) We propose a 4D Mask Fusion module, integrating memory-conditioned attention and logits equalization for robust embedding fusion. Extensive experiments demonstrate the effectiveness and flexibility of the proposed method on 4D human-centric parsing tasks, achieving up to 93.3% acceleration compared to the previous state-of-the-art method, which was limited to parsing fixed classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。