解决图像压缩修复中保真度、人眼感知与机器偏好难兼顾的问题
FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration

- 两阶段设计:先用流匹配恢复语义结构,再确定性还原高频纹理
- 在JPEG压缩下实现8.3%的峰值信噪比提升,同时保持感知质量
- 适合需要高保真且避免生成幻觉的图像修复任务
图像修复质量可从像素级保真度、人类感知和下游机器偏好三个维度评估。现有方法通常只优化其中一项:保真度导向模型趋向于条件均值,导致输出过平滑;生成式方法会幻化出看似合理但事实错误的纹理,损害真实保真度和下游任务准确率。为应对这一三重权衡,我们提出FDIR,一种两阶段架构,通过互补归纳偏置解耦冲突需求:质量引导的一步流匹配(QO-Flow)在隐空间单次前向传播中恢复全局语义结构;流条件的细节精修(FCDR)在像素空间确定性地还原高频纹理并抑制生成幻觉。大量实验表明,FDIR在保真度上表现更优,兼具良好的感知-保真平衡,并在机器偏好上达到竞争力。
原文摘要 · Abstract (English)
Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference. However, existing lossy compression restoration methods optimize for at most one of these criteria: fidelity-oriented models often regress toward conditional means and produce over-smoothed outputs, while generative approaches hallucinate plausible but factually incorrect textures that degrade both ground-truth fidelity and downstream task accuracy. To navigate this three-way tradeoff, we propose FDIR, a two-stage architecture that decouples the conflicting demands through complementary inductive biases: Quality-Guided One-Step Flow Matching (QO-Flow) recovers global semantic structure in latent space via a single forward pass, while Flow-Conditioned Detail Refinement (FCDR) deterministically restores high-frequency textures and suppresses generative hallucinations in pixel space. Extensive experiments demonstrate that FDIR achieves superior fidelity, with a favorable perceptual-fidelity balance and competitive machine preference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。