用稀疏关键点实现超低码率人体视频压缩与精确顶点预测
Sparse2Dense: A Keypoint-driven Generative Framework for Human Video Compression and Vertex Prediction
- 以极稀疏3D关键点为编码符号,驱动生成式视频重建
- 在低于100kbps码率下保持视频清晰度与运动连贯性
- 适合实时人体动作分析、虚拟人动画等低带宽场景
针对带宽受限的多媒体应用,同时实现超低码率人体视频压缩与精确顶点预测仍是一大挑战,这需要动态运动建模、细节外观合成与几何一致性之间的协调。为此,我们提出Sparse2Dense,一种基于关键点的生成框架,利用极稀疏的3D关键点作为紧凑传输符号,实现超低码率人体视频压缩与精准的人体顶点预测。其核心创新在于多任务学习与关键点感知的深度生成模型,能够通过紧凑的3D关键点编码复杂人体运动,并利用这些稀疏关键点估计稠密运动以实现具有时间一致性和真实纹理的视频合成。此外,集成的顶点预测器通过与视频生成联合优化,学习人体顶点几何结构,确保视觉内容与几何结构的一致性。大量实验表明,所提出的Sparse2Dense框架在传统/生成式视频编码器中均实现了具有竞争力的压缩性能,同时支持下游几何应用中的精确人体顶点预测。因此,Sparse2Dense有望推动带宽高效的以人为中心媒体传输,如实时动作分析、虚拟人动画和沉浸式娱乐。
原文摘要 · Abstract (English)
For bandwidth-constrained multimedia applications, simultaneously achieving ultra-low bitrate human video compression and accurate vertex prediction remains a critical challenge, as it demands the harmonization of dynamic motion modeling, detailed appearance synthesis, and geometric consistency. To address this challenge, we propose Sparse2Dense, a keypoint-driven generative framework that leverages extremely sparse 3D keypoints as compact transmitted symbols to enable ultra-low bitrate human video compression and precise human vertex prediction. The key innovation is the multi-task learning-based and keypoint-aware deep generative model, which could encode complex human motion via compact 3D keypoints and leverage these sparse keypoints to estimate dense motion for video synthesis with temporal coherence and realistic textures. Additionally, a vertex predictor is integrated to learn human vertex geometry through joint optimization with video generation, ensuring alignment between visual content and geometric structure. Extensive experiments demonstrate that the proposed Sparse2Dense framework achieves competitive compression performance for human video over traditional/generative video codecs, whilst enabling precise human vertex prediction for downstream geometry applications. As such, Sparse2Dense is expected to facilitate bandwidth-efficient human-centric media transmission, such as real-time motion analysis, virtual human animation, and immersive entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。