用3D人体姿态融合视觉模型,提升复杂场景下的动作识别鲁棒性
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
- 结合上下文动态与显式人体姿态,实现空间感知的动作理解
- 在复杂遮挡场景下,性能超越三种基线模型
- 适合需要精准动作识别的机器人交互场景
为使具身智能体有效理解并互动于周围世界,需具备基于物理空间的人类动作理解能力。现有动作识别模型多依赖RGB视频,仅学习表面模式与动作标签间的相关性,在复杂场景中难以捕捉真实物理交互动态与人体姿态。本文提出一种新架构,通过融合V-JEPA 2的上下文预测世界动态与CoMotion的显式、抗遮挡人体姿态数据,实现动作识别的空间化。模型在InHARD(通用动作识别)和UCF-19-Y-OCC(高遮挡动作识别)基准上验证,显著优于三种基线,尤其在复杂遮挡场景中表现突出。结果表明,动作识别应依赖空间理解而非统计模式匹配。
原文摘要 · Abstract (English)
For embodied agents to effectively understand and interact within the world around them, they require a nuanced comprehension of human actions grounded in physical space. Current action recognition models, often relying on RGB video, learn superficial correlations between patterns and action labels, so they struggle to capture underlying physical interaction dynamics and human poses in complex scenes. We propose a model architecture that grounds action recognition in physical space by fusing two powerful, complementary representations: V-JEPA 2's contextual, predictive world dynamics and CoMotion's explicit, occlusion-tolerant human pose data. Our model is validated on both the InHARD and UCF-19-Y-OCC benchmarks for general action recognition and high-occlusion action recognition, respectively. Our model outperforms three other baselines, especially within complex, occlusive scenes. Our findings emphasize a need for action recognition to be supported by spatial understanding instead of statistical pattern recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。