用户输入文字即可自定义图像重点区域,灵活调节画质分配。
Customizable ROI-Based Deep Image Compression
- 用文本控制生成感兴趣区域掩码,支持自由定义重点
- 可调非重点区域压缩程度,实现重点与背景画质平衡
- 在潜空间融合掩码与率失真优化先验,提升重建质量
基于感兴趣区域(ROI)的图像压缩通过优先分配比特提升重点区域的重建质量。然而,随着用户需求日益多样化,现有方法因预设固定ROI且无法灵活调节重点与非重点区域之间的质量权衡,难以满足个性化需求。本文提出一种可定制化的深度图像压缩范式:首先设计文本控制掩码获取(TMA)模块,用户仅需输入语义文本即可自定义ROI;其次提出可定制值分配(CVA)机制,允许用户动态调整非重点区域的压缩强度以管理质量平衡;最后引入潜空间掩码注意力(LMA)模块,在潜空间融合掩码空间先验与率失真优化先验,优化源图像的潜表示。实验表明,该方法有效支持用户对ROI定义与获取的定制化需求,并实现对重点与非重点区域重建质量的灵活调控。
原文摘要 · Abstract (English)
Region of Interest (ROI)-based image compression optimizes bit allocation by prioritizing ROI for higher-quality reconstruction. However, as the users (including human clients and downstream machine tasks) become more diverse, ROI-based image compression needs to be customizable to support various preferences. For example, different users may define distinct ROI or require different quality trade-offs between ROI and non-ROI. Existing ROI-based image compression schemes predefine the ROI, making it unchangeable, and lack effective mechanisms to balance reconstruction quality between ROI and non-ROI. This work proposes a paradigm for customizable ROI-based deep image compression. First, we develop a Text-controlled Mask Acquisition (TMA) module, which allows users to easily customize their ROI for compression by just inputting the corresponding semantic \emph{text}. It makes the encoder controlled by text. Second, we design a Customizable Value Assign (CVA) mechanism, which masks the non-ROI with a changeable extent decided by users instead of a constant one to manage the reconstruction quality trade-off between ROI and non-ROI. Finally, we present a Latent Mask Attention (LMA) module, where the latent spatial prior of the mask and the latent Rate-Distortion Optimization (RDO) prior of the image are extracted and fused in the latent space, and further used to optimize the latent representation of the source image. Experimental results demonstrate that our proposed customizable ROI-based deep image compression paradigm effectively addresses the needs of customization for ROI definition and mask acquisition as well as the reconstruction quality trade-off management between the ROI and non-ROI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。