arXiv:2510.15347eess.IVcs.MM2025-10中稿 · IEEE Transactions …被引 1

让视频编码直接对齐机器视觉需求,显著降低传输码率。

Symmetric Entropy-Constrained Video Coding for Machines

  • 构建编码器与视觉模型的对称对齐,指导编码保留语义、舍弃无关信息。
  • 在实例分割等任务上比H.266/VVC节省超97%码率,最高达97.6%。
  • 适合多任务机器视觉场景,无需针对特定模型重训练。

随着视频传输越来越多服务于机器视觉系统(MVS)而非人类视觉系统(HVS),面向机器的视频编码(VCM)成为关键研究方向。现有VCM方法常将编解码器绑定于特定下游模型,需重新训练或监督数据,限制了多任务泛化能力。近期统一的VCM框架采用视觉骨干(VB)和视觉基础模型(VFM)支持多种视频理解任务,但主要通过保持语义一致性或抑制非语义信息,未深入探索如何在VB/VFM指导下直接关联视频编码与理解。为此,本文提出面向机器的对称熵约束视频编码框架(SEC-VCM),建立编解码器与VB之间的对称对齐,使编码器可利用VB的表征能力保留语义并剔除对MVS无用信息。具体地,双向熵约束(BiEC)机制通过抑制条件熵,在解码与VB编码间实现对称性,明确处理有益于机器理解的信息并压缩冗余内容。此外,语义-像素双路径融合(SPDF)模块注入像素级先验,通过融合提升机器导向的重建质量,抑制有害伪影。在经典视频理解任务及基于多模态大模型(MLLM)的任务上,实验结果表明其达到最先进的率-任务性能,在视频实例分割(37.4%)、视频目标分割(29.8%)、目标检测(46.2%)、多目标跟踪(44.9%)和基于MLLM的视频定位(97.6%)上相较H.266/VVC参考软件VTM实现显著码率节省。

原文摘要 · Abstract (English)

As video transmission increasingly serves machine vision systems (MVS) instead of human vision systems (HVS), video coding for machines (VCM) has become a critical research topic. Existing VCM methods often bind codecs to specific downstream models, requiring retraining or supervised data, thus limiting generalization in multi-task scenarios. Recently, unified VCM frameworks have employed visual backbones (VB) and visual foundation models (VFM) to support multiple video understanding tasks with a single codec. They mainly utilize VB/VFM to maintain semantic consistency or suppress non-semantic information, but seldom explore how to directly link video coding with understanding under VB/VFM guidance. Hence, we propose a Symmetric Entropy-Constrained Video Coding framework for Machines (SEC-VCM). It establishes a symmetric alignment between the video codec and VB, allowing the codec to leverage VB's representation capabilities to preserve semantics and discard MVS-irrelevant information. Specifically, a bi-directional entropy-constraint (BiEC) mechanism ensures symmetry between the process of video decoding and VB encoding by suppressing conditional entropy. This helps the codec to explicitly handle semantic information beneficial to MVS while squeezing useless information. Furthermore, a semantic-pixel dual-path fusion (SPDF) module injects pixel-level priors into the final reconstruction. Through semantic-pixel fusion, it suppresses artifacts harmful to MVS and improves machine-oriented reconstruction quality. Experimental results on classical video understanding tasks and MLLM-based tasks show SOTA rate-task performance. It achieves significant bitrate savings over H.266/VVC reference software VTM on video instance segmentation (37.4%), video object segmentation (29.8%), object detection (46.2%), multiple object tracking (44.9%), and MLLM-based video grounding (97.6%).

视频编码机器视觉熵约束多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。