arXiv:2607.19515eess.IVcs.CV2026-07

BLUE压缩视频在不损失语义理解能力的前提下,大幅降低推理开销。

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

论文配图:BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics
图 1 · 摘自论文原文
  • 通过去除静态背景冗余,保留关键活动信息进行压缩
  • 在VIRAT和CHAD数据集上,语义评分几乎无下降(差值≈-0.01)
  • 适合需要降低计算成本的视觉语言模型监控系统使用

持续监控视频给企业级视觉分析系统带来日益增长的存储、传输和推理负担。尽管现代编码器如H.265能降低人类可观看视频的码率,但激进压缩可能损害下游计算机视觉性能,且未必减少用于语义理解的视觉语言模型(VLM)推理调用次数。本文评估了BLUE——一种针对固定摄像头监控场景的压缩方法,其通过抑制静态背景冗余并保留前景活动,对基于VLM的事件与异常理解的影响。在两个监控数据集VIRAT(含227个配对事件样本,来自106段视频)和CHAD(含54段人行动异常视频)上,采用相同帧索引对比原始H.265与BLUE压缩后的H.265视频,并通过盲评协议以注释为基准评估VLM生成描述的质量。结果表明,语义推断质量无显著下降:在VIRAT上,原始与BLUE压缩的平均得分差异约为-0.01(0-10分制);在CHAD上,原始与BLUE得分分别为4.31与4.26。压缩率与语义得分变化无关(VIRAT上相关系数r=0.004),说明更高压缩并不预示语义质量损失。此外,蓝色压缩将CHAD中跳过率高的P帧比例从1.4%提升至53.2%,理论上可实现53%的VLM调用减少。这些发现表明,BLUE可作为面向机器的监控视频压缩层,在降低带宽与推理成本的同时,保持VLM语义性能。

原文摘要 · Abstract (English)

Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise video analytics systems. While modern codecs such as H.265 reduce bitrate for human-viewable video, aggressive compression can degrade downstream computer-vision performance and does not necessarily reduce the number of vision-language model (VLM) inference calls required for semantic video understanding. This paper evaluates BLUE, a fixed-camera surveillance compression approach that suppresses static-background redundancy while preserving foreground activity, for its effect on VLM-based event and anomaly understanding. We compare raw H.265 and BLUE-compressed H.265 video on two surveillance datasets: VIRAT, comprising 227 paired event samples from 106 clips, and CHAD, comprising 54 human-activity anomaly clips. For each pair, the same frame index is evaluated using a VLM captioning pipeline, and outputs are scored against annotation-derived ground truth using a blind judging protocol. The results show no measurable degradation in semantic inference quality. On VIRAT, the mean VLM score remains effectively unchanged between raw H.265 and BLUE, with a mean difference of approximately -0.01 on a 0-10 scale. On CHAD, raw H.265 and BLUE obtain near-equivalent mean scores of 4.31 and 4.26, respectively. Compression saving is also uncorrelated with VLM score change on VIRAT (r = 0.004), indicating that higher BLUE compression does not predict semantic quality loss. Beyond storage reduction, BLUE increases the share of skip-heavy P-frames on CHAD from 1.4% to 53.2%, enabling an estimated 53% reduction in VLM calls through packet-size-based frame skipping. These findings suggest that BLUE functions as a machine-centric compression layer for surveillance video, reducing bandwidth and inference cost while preserving VLM semantic performance.

视频压缩视觉语言模型监控分析机器感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。