用视频训练的视觉模型,能感知人体状态并预测未来动作。
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

- 通过视频锚定预测,防止密集感知失效。
- 参数减少2.7倍,仍优于专业模型在姿态与重识别任务表现。
- 首个不退化的预测头,兼顾感知与预测能力。
能够理解人类的机器应能感知当下并预测未来。现有以人为主体的视觉模型虽在静态密集感知上达到顶尖水平,但难以处理运动与预测。本文提出 Human-JEPA,一种基于视频训练的人类中心视觉模型,采用锚定预测机制:将密集目标固定在初始化的冻结副本上,防止密集感知的隐性崩溃;同时用纯时序前后分割替代块掩码,避免五点动作税与十七点重识别崩溃。在冻结探测器下,Human-JEPA 在姿态估计与行人重识别任务上仅需 2.7 倍更少的参数便超越像素锚定的专业模型,牺牲高分辨率密集分割能力,其发布的预测头是首个不会退化的前瞻模型。单一安全适配模型即可同时实现对人类的感知与预测。
原文摘要 · Abstract (English)
Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。