arXiv:2512.15262eess.IVcs.MM2025-12中稿 · as a PAPER and for…

利用音视频同步性实现高效联合压缩,显著降低人脸视频传输码率。

Audio-Visual Cross-Modal Compression for Generative Face Video Coding

  • 通过统一扩散模型对音视频特征进行跨模态对齐与共享编码
  • 在极低码率下仍能实现音视频的高质量同步重建,性能超越VVC标准
  • 适合需要低带宽高保真音视频通信的场景,如远程会议

生成式人脸视频编码(GFVC)在视频会议等现代应用中至关重要,但现有方法主要关注视频运动而忽视了音频带来的显著码率开销。尽管音频与口型动作之间存在明确关联,这一跨模态一致性尚未被系统用于压缩。为此,我们提出音频-视觉跨模态压缩(AVCC)框架,联合压缩音频与视频流。该框架从视频中提取运动信息,对音频特征进行分词,并通过统一的音视频扩散过程实现对齐,使两者可从共享表示中同步重建。在极端低码率场景下,甚至可实现单模态从另一模态的重构。实验表明,AVCC在码率-失真性能上显著优于通用视频编码(VVC)标准及当前最优的GFVC方案,为更高效的多模态通信系统铺平道路。

原文摘要 · Abstract (English)

Generative face video coding (GFVC) is vital for modern applications like video conferencing, yet existing methods primarily focus on video motion while neglecting the significant bitrate contribution of audio. Despite the well-established correlation between audio and lip movements, this cross-modal coherence has not been systematically exploited for compression. To address this, we propose an Audio-Visual Cross-Modal Compression (AVCC) framework that jointly compresses audio and video streams. Our framework extracts motion information from video and tokenizes audio features, then aligns them through a unified audio-video diffusion process. This allows synchronized reconstruction of both modalities from a shared representation. In extremely low-rate scenarios, AVCC can even reconstruct one modality from the other. Experiments show that AVCC significantly outperforms the Versatile Video Coding (VVC) standard and state-of-the-art GFVC schemes in rate-distortion performance, paving the way for more efficient multimodal communication systems.

音视频压缩生成式编码跨模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。