用稳定扩散先验提升人脸压缩的视觉与语义一致性。
FaSDiff: Balancing Perception and Semantics in Face Compression via Stable Diffusion Priors
- 引入高频敏感压缩器捕捉细节,生成引导扩散模型的视觉提示。
- 设计混合低频增强模块,保留语义结构,避免重建失真。
- 兼顾人眼感知与机器识别,适合需要高保真人脸数据的应用。
随着人脸图像在各类应用中日益普及,针对面部语义的高效压缩对存储和传输至关重要。现有基于学习的人脸图像压缩方法在低比特率下常出现重建质量下降的问题。直接应用基于扩散的生成先验会导致下游机器视觉任务性能不佳,主要因为高频细节保留不足。本文提出 FaSDiff(面部图像压缩的稳定扩散先验),一种新型扩散驱动的压缩框架,旨在同时提升视觉保真度与语义一致性。FaSDiff 引入高频敏感压缩器以捕捉细微特征,并生成鲁棒的视觉提示,引导扩散模型重建。为缓解低频失真,进一步设计混合低频增强模块,实现语义结构的解耦与保留,确保扩散先验在重建过程中的稳定调制。通过联合优化感知质量与语义保持,FaSDiff 有效平衡了人类视觉感知与机器视觉准确性。大量实验表明,其在感知指标与下游任务表现上均优于当前最优方法。
原文摘要 · Abstract (English)
With the increasing deployment of facial image data across a wide range of applications, efficient compression tailored to facial semantics has become critical for both storage and transmission. While recent learning-based face image compression methods have achieved promising results, they often suffer from degraded reconstruction quality at low bit rates. Directly applying diffusion-based generative priors to this task leads to suboptimal performance in downstream machine vision tasks, primarily due to poor preservation of high-frequency details. In this work, we propose FaSDiff (\textbf{Fa}cial Image Compression with a \textbf{S}table \textbf{Diff}usion Prior), a novel diffusion-driven compression framework designed to enhance both visual fidelity and semantic consistency. FaSDiff incorporates a high-frequency-sensitive compressor to capture fine-grained details and generate robust visual prompts for guiding the diffusion model. To address low-frequency degradation, we further introduce a hybrid low-frequency enhancement module that disentangles and preserves semantic structures, enabling stable modulation of the diffusion prior during reconstruction. By jointly optimizing perceptual quality and semantic preservation, FaSDiff effectively balances human visual fidelity and machine vision accuracy. Extensive experiments demonstrate that FaSDiff outperforms state-of-the-art methods in both perceptual metrics and downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。