arXiv:2505.16152eess.IVcs.CV2025-05被引 4

用可交互语义压缩人体视频,实现低码率下高质量重建与可控编辑。

Compressing Human Body Video with Interactive Semantics: A Generative Approach

  • 通过3D人体模型分离运动信号,生成可配置的语义嵌入。
  • 在极低码率下优于VVC和最新生成式压缩方案。
  • 支持直接交互编辑,适合元宇宙数字人通信场景。

本文提出一种基于交互语义的人体视频压缩方法,使视频编码具备交互性和可控性,可通过解码比特流中的语义级表示进行操控。具体而言,编码器利用3D人体模型将人体信号的非线性动态和复杂运动分解为一系列可配置的嵌入表示,这些嵌入可被可控编辑、紧凑压缩并高效传输。解码器则从解码后的语义中演化出基于网格的运动场,实现高质量人体视频重建。实验表明,该框架在超低码率范围内,相较于最先进的视频编码标准Versatile Video Coding(VVC)及最新的生成式压缩方案,均展现出优异的压缩性能。此外,该框架无需额外预/后处理即可实现交互式人体视频编码,有望推动未来元宇宙中数字人通信的发展。

原文摘要 · Abstract (English)

In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body signal into a series of configurable embeddings, which are controllably edited, compactly compressed, and efficiently transmitted. Moreover, the proposed decoder can evolve the mesh-based motion fields from these decoded semantics to realize the high-quality human body video reconstruction. Experimental results illustrate that the proposed framework can achieve promising compression performance for human body videos at ultra-low bitrate ranges compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes. Furthermore, the proposed framework enables interactive human body video coding without any additional pre-/post-manipulation processes, which is expected to shed light on metaverse-related digital human communication in the future.

人体视频生成压缩交互编码元宇宙

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。