模仿人类视觉与推理机制,提升人体运动预测精度。
HVIS: A Human-like Vision and Inference System for Human Motion Prediction
- 模拟人眼视觉与大脑推理,分阶段处理运动信息。
- 在Human3.6M等数据集上提升11.1%至19.8%性能。
- 适合需要高精度运动预测的机器人与动画应用。
理解人体运动中的时空依赖与多尺度效应,对运动预测至关重要。尽管人类天生具备此类能力,机器却难以模拟。为此,我们提出人类类视觉与推理系统(HVIS),用于人体运动预测,旨在模仿人类的观察与预测过程。HVIS包含两个模块:类人视觉编码(HVE)与类人运动推理(HMI)。HVE模块模拟并优化人类视觉过程,引入视网膜类组件分别捕捉时空信息,避免冗余干扰;同时设计视觉皮层类组件,分层提取复杂运动特征,关注人体姿态的全局与局部信息。HMI模块模拟人脑多阶段学习模型:自发学习网络模拟神经元生成过程,实现未来运动的对抗性生成;刻意学习网络则针对难训练关节进行优化,防止误导性学习。实验表明,本方法在Human3.6M、CMU Mocap和G3D数据集上分别显著优于现有方法19.8%、15.7%和11.1%,达到新最先进水平。
原文摘要 · Abstract (English)
Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it proves to be markedly more challenging for machines to emulate. To bridge the gap, we propose the Human-like Vision and Inference System (HVIS) for human motion prediction, which is designed to emulate human observation and forecast future movements. HVIS comprises two components: the human-like vision encode (HVE) module and the human-like motion inference (HMI) module. The HVE module mimics and refines the human visual process, incorporating a retina-analog component that captures spatiotemporal information separately to avoid unnecessary crosstalk. Additionally, a visual cortex-analogy component is designed to hierarchically extract and treat complex motion features, focusing on both global and local features of human poses. The HMI is employed to simulate the multi-stage learning model of the human brain. The spontaneous learning network simulates the neuronal fracture generation process for the adversarial generation of future motions. Subsequently, the deliberate learning network is optimized for hard-to-train joints to prevent misleading learning. Experimental results demonstrate that our method achieves new state-of-the-art performance, significantly outperforming existing methods by 19.8% on Human3.6M, 15.7% on CMU Mocap, and 11.1% on G3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。