arXiv:2507.01608cs.CVeess.IV2025-07被引 3

让压缩域特征更有语义,用小模型适配实现高效视觉推理

Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference

  • 针对压缩域特征语义弱问题,设计感知导向的隐空间编码方法
  • 仅需轻量适配器微调,参数量远低于全模型微调
  • 在保持压缩效率的同时显著提升图像理解任务性能

近年来,压缩域语义推理主要依赖以均方误差(MSE)优化的图像编码模型,但此类优化导致隐空间语义信息贫乏,限制了下游任务表现。且高精度推理常需微调整个视觉模型,计算开销大,尤其对大模型而言。为此,本文提出感知导向隐编码(POLC),通过增强隐空间的语义丰富性,实现高性能压缩域语义推理。得益于语义丰富的隐表示,POLC仅需一个即插即用的适配器进行微调,大幅降低参数量。实验表明,POLC在速率-感知性能上媲美前沿生成式图像编码方法,同时在各类视觉任务中表现显著提升,且微调开销极低。

原文摘要 · Abstract (English)

In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization tends to yield latent spaces with limited semantic richness, which hinders effective semantic inference in downstream tasks. Moreover, achieving high performance with these models often requires fine-tuning the entire vision model, which is computationally intensive, especially for large models. To address these problems, we introduce Perception-Oriented Latent Coding (POLC), an approach that enriches the semantic content of latent features for high-performance compressed domain semantic inference. With the semantically rich latent space, POLC requires only a plug-and-play adapter for fine-tuning, significantly reducing the parameter count compared to previous MSE-oriented methods. Experimental results demonstrate that POLC achieves rate-perception performance comparable to state-of-the-art generative image coding methods while markedly enhancing performance in vision tasks, with minimal fine-tuning overhead. Code is available at https://github.com/NJUVISION/POLC.

压缩域推理隐空间编码轻量微调语义增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。