构建合成数据提升视频人体密集预测的时序一致性。
From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction
- 用可扩展合成数据生成带精确几何标签的运动对齐视频序列。
- 在THuman2.1和Hi4D上达到当前最佳性能,且能泛化到真实视频。
- 适合做视频人体理解、动作分析与三维重建的研究者参考。
本文聚焦于视频序列中人体中心密集预测的时序一致性挑战。现有模型虽单帧精度高,但在运动、遮挡和光照变化下常出现闪烁,且极少有支持多任务的成对人体视频监督。为此,我们提出一种可扩展的合成数据流水线,生成具有像素级深度、法向量和掩码的逼真人体图像与运动对齐序列。与以往静态合成方法不同,该流水线同时提供帧级标签用于空间学习和序列级监督用于时序学习。基于此,我们训练了一个统一的ViT密集预测器,通过CSE嵌入显式注入人体几何先验,并在特征融合后使用轻量级通道重加权模块提升几何特征可靠性。采用两阶段训练策略——先静态预训练获取鲁棒空间表征,再通过动态序列监督优化时序一致性。大量实验表明,模型在THuman2.1和Hi4D上表现领先,且能有效泛化至真实场景视频。
原文摘要 · Abstract (English)
In this work, we focus on the challenge of temporally consistent human-centric dense prediction across video sequences. Existing models achieve strong per-frame accuracy but often flicker under motion, occlusion, and lighting changes, and they rarely have paired human video supervision for multiple dense tasks. We address this gap with a scalable synthetic data pipeline that generates photorealistic human frames and motion-aligned sequences with pixel-accurate depth, normals, and masks. Unlike prior static data synthetic pipelines, our pipeline provides both frame-level labels for spatial learning and sequence-level supervision for temporal learning. Building on this, we train a unified ViT-based dense predictor that (i) injects an explicit human geometric prior via CSE embeddings and (ii) improves geometry-feature reliability with a lightweight channel reweighting module after feature fusion. Our two-stage training strategy, combining static pretraining with dynamic sequence supervision, enables the model first to acquire robust spatial representations and then refine temporal consistency across motion-aligned sequences. Extensive experiments show that we achieve state-of-the-art performance on THuman2.1 and Hi4D and generalize effectively to in-the-wild videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。