ReLViC通过分散编码块和可控依赖,提升视频传输在丢包下的稳定性。
ReLViC: Loss-Resilient Learned Video Coding with Dispersed Packetization and Controllable Packet Dependencies

- 将相邻的隐变量分散到不同数据包,降低丢包影响范围。
- 在严重丢包下重建质量优于H.265+FEC和GRACE基准模型。
- 无需重训练即可调节压缩效率与错误传播的平衡,适合实时视频传输。
丢包会严重损害学习型视频编码性能,因为缺失的隐变量会影响空间重建和时间预测。本文提出ReLViC,一种联合处理隐变量编码与丢包恢复的鲁棒性学习视频编码框架。ReLViC将空间相邻的隐变量分散至不同数据包,并采用双用途Transformer,在编码时估计熵模型参数,在接收端重构丢失的隐变量。通过周期性重置的数据包上下文拓扑结构,由片段长度参数化,可调控压缩效率与错误传播范围,且无需重新训练。采用三阶段渐进式训练:先建立单帧编码,再学习时间上下文用于熵建模,最后在模拟丢包条件下优化掩码隐变量的恢复能力。基于突发丢包测试序列的实验表明,ReLViC在严重丢包场景下重建更稳定,性能优于采用前向纠错(FEC)保护的H.265及另一款鲁棒学习视频编解码器GRACE。
原文摘要 · Abstract (English)
Packet loss can severely impair learned video coding because missing latent tokens compromise both spatial reconstruction and temporal prediction. We present ReLViC, a loss-resilient learned video coding framework that jointly addresses latent coding and packet-loss recovery. ReLViC disperses spatially adjacent latent tokens across packets and employs a dual-purpose Transformer to estimate entropy-model parameters during coding and reconstruct missing latent tokens at the receiver. It controls packet dependencies through a periodic-reset packet-context topology parameterized by the segment length, thereby tuning the trade-off between compression efficiency and error-propagation range without retraining. A three-stage progressive training procedure establishes single-frame coding, learns temporal context for entropy modeling, and then optimizes the recovery of masked latent tokens under simulated packet loss. Experiments using burst-loss traces evaluate ReLViC against H.265 protected by Reed--Solomon forward error correction (FEC) and GRACE, a loss-resilient learned video codec. ReLViC delivers more stable reconstruction and outperforms both baselines under severe packet loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。