arXiv:2603.03615cs.CV2026-03中稿 · CVPR

提出新型注意力机制,让多视角图像压缩更高效。

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

  • 用全景视差注意力显式建模不同视角间的相关性
  • 在6个视角下比特率降低最高达24.18%
  • 适合需要高效率多视角视频压缩的场景

多视角图像压缩(MIC)通过利用图像间相关性实现高效压缩,在3D应用中至关重要。分布式多视角图像压缩(DMIC)在无需编码端传输视角间信息的前提下,性能接近传统MIC。然而现有方法对所有视角一视同仁,忽略其相关性差异,导致编码效率受限。为此,本文提出一种通用的全景视差注意力机制(OPAM),可显式建模任意信息源之间的相关性和特征对齐。基于此,设计了视差多信息融合模块(PMIFM),并将其融入联合解码器与熵模型,构建端到端的DMIC框架ParaHydra。大量实验表明,ParaHydra是首个显著超越现有MIC编解码器的DMIC方法,计算开销低。随着输入视角数增加,性能优势更明显:相比LDMIC,WildTrack(3)上节省19.72%比特率,WildTrack(6)上最高达24.18%,编码效率提升最高达65倍(解码)和34倍(编码)。

原文摘要 · Abstract (English)

Multi-view image compression (MIC) aims to achieve high compression efficiency by exploiting inter-image correlations, playing a crucial role in 3D applications. As a subfield of MIC, distributed multi-view image compression (DMIC) offers performance comparable to MIC while eliminating the need for inter-view information at the encoder side. However, existing methods in DMIC typically treat all images equally, overlooking the varying degrees of correlation between different views during decoding, which leads to suboptimal coding performance. To address this limitation, we propose a novel $\textbf{OmniParallax Attention Mechanism}$ (OPAM), which is a general mechanism for explicitly modeling correlations and aligned features between arbitrary pairs of information sources. Building upon OPAM, we propose a Parallax Multi Information Fusion Module (PMIFM) to adaptively integrate information from different sources. PMIFM is incorporated into both the joint decoder and the entropy model to construct our end-to-end DMIC framework, $\textbf{ParaHydra}$. Extensive experiments demonstrate that $\textbf{ParaHydra}$ is $\textbf{the first DMIC method}$ to significantly surpass state-of-the-art MIC codecs, while maintaining low computational overhead. Performance gains become more pronounced as the number of input views increases. Compared with LDMIC, $\textbf{ParaHydra}$ achieves bitrate savings of $\textbf{19.72%}$ on WildTrack(3) and up to $\textbf{24.18%}$ on WildTrack(6), while significantly improving coding efficiency (as much as $\textbf{65}\times$ in decoding and $\textbf{34}\times$ in encoding).

多视角压缩注意力机制图像编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。