arXiv:2608.21837cs.CVcs.MM2026-08

用损坏码流学语义,提升恶劣视频理解能力

Towards Bitstream-corrupted Harsh Visual Understanding: Through Bitstream Language Modeling as Robust Semantic Priors

论文配图:Towards Bitstream-corrupted Harsh Visual Understanding: Through Bitstream Language Modeling as Robust Semantic Priors
图 1 · 摘自论文原文
  • 通过码流语言建模提取字节级语义线索
  • 在视频修复/字幕/姿态估计上分别提升2.51dB、0.20、0.18
  • 适合做鲁棒视频理解的工程师和研究者

码流损坏的恶劣视觉理解(BcHVU)旨在理解由严重损坏码流解码出的劣质视频,现有视觉模型难以应对。本文提出比特流语言建模作为鲁棒语义先验(BLMSP),通过字节级建模与跨编码器语义蒸馏,从多种损坏码流中提取鲁棒语义,并注入通用视觉模型。构建了包含60.7万段损坏码流和28.7万对劣质视频的大型多源数据集CHP。实验表明,该方法在视频修复、字幕生成和人体姿态估计上分别提升2.51 dB(PSNR)、0.20(CIDEr)和0.18([email protected])。结果证明,损坏码流可作为解决像素失真与语义丢失的有效语义先验。

原文摘要 · Abstract (English)

Bitstream-corrupted Harsh Visual Understanding (BcHVU) aims to understand harshly degraded videos originally decoded from a severely corrupted bitstream in real-world multimedia communication. The ill-posed nature of BcHVU poses a major challenge for existing vision models, as even subtle bitstream corruption can lead to irreversible pixel distortion and significant semantic loss. To address these challenges in BcHVU, we propose Bitstream Language Modeling as Robust Semantic Priors (BLMSP), a framework for learning and injecting bitstream-native semantic cues. Our proposed BLMSP framework learns to extract bitstream-native semantic cues by bitstream language modeling, and leverages them as priors by injecting into off-the-shelf vision models of BcHVU tasks. Specifically, we present a Video Bitstream Byte Model (VBBM) that integrates byte-level modeling and cross-codec semantic distillation, enabling it to interpret robust semantics from byte sequences in multiple corrupted bitstream formats. The learned bitstream semantics are leveraged as robust priors and fused into BcHVU model backbones for improving the quality of video restoration, captioning, and human pose estimation. To train BLMSP, we construct a large-scale multi-source Corrupted-bitstream Harsh-video Paired (CHP) dataset containing 607k corrupted bitstream segments and 287k paired harsh video clips. Extensive experimental results show that the learned bitstream priors improve video restoration, captioning, and human pose estimation by 2.51 dB in PSNR, 0.20 in CIDEr, and 0.18 in [email protected] on average, respectively. These results demonstrate that corrupted bitstream can serve as robust semantic priors in solving pixel distortion and semantic loss in BcHVU.

视频理解码流鲁棒性语义先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。