arXiv:2501.04329cs.CV2025-01被引 11

新压缩方法兼顾人眼与机器视觉,提升多任务性能。

An Efficient Adaptive Compression Method for Human Perception and Machine Vision Tasks

  • 自适应选择特征子集,平衡人眼与机器任务优化。
  • 用轻量微调策略提升分割、检测等任务表现。
  • 可无缝接入现有压缩模型,适合多场景应用。

现有神经图像/视频压缩方法主要优化人眼感知,但随着人工智能发展,图像视频大量用于机器视觉任务,传统方法表现不佳。本文提出高效自适应压缩(EAC)方法,包含两个模块:一是自适应压缩机制,从潜在特征中动态选择子集,兼顾人眼感知与多种机器视觉任务(如分割、检测);二是任务特定适配器,采用参数高效的增量微调策略,激发下游分析网络在特定任务上的性能。该方法在保持人眼视觉质量的同时,显著提升多个机器视觉任务表现。实验在VOC2007、ILSVRC2012、VOC2012、COCO、UCF101和DAVIS等数据集上验证,结果表明其能有效集成至Ballé2018、Cheng2020、DVC和FVC等主流压缩方法中。

原文摘要 · Abstract (English)

While most existing neural image compression (NIC) and neural video compression (NVC) methodologies have achieved remarkable success, their optimization is primarily focused on human visual perception. However, with the rapid development of artificial intelligence, many images and videos will be used for various machine vision tasks. Consequently, such existing compression methodologies cannot achieve competitive performance in machine vision. In this work, we introduce an efficient adaptive compression (EAC) method tailored for both human perception and multiple machine vision tasks. Our method involves two key modules: 1), an adaptive compression mechanism, that adaptively selects several subsets from latent features to balance the optimizations for multiple machine vision tasks (e.g., segmentation, and detection) and human vision. 2), a task-specific adapter, that uses the parameter-efficient delta-tuning strategy to stimulate the comprehensive downstream analytical networks for specific machine vision tasks. By using the above two modules, we can optimize the bit-rate costs and improve machine vision performance. In general, our proposed EAC can seamlessly integrate with existing NIC (i.e., Ballé2018, and Cheng2020) and NVC (i.e., DVC, and FVC) methods. Extensive evaluation on various benchmark datasets (i.e., VOC2007, ILSVRC2012, VOC2012, COCO, UCF101, and DAVIS) shows that our method enhances performance for multiple machine vision tasks while maintaining the quality of human vision.

图像压缩机器视觉自适应轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。