arXiv:2602.21402cs.CV2026-02被引 1

FlowFixer修复主体生成中的细节丢失问题,提升图像保真度。

FlowFixer: Towards Detail-Preserving Subject-Driven Generation

  • 通过图像到图像的直接转换,避免语言提示带来的歧义。
  • 自监督训练数据模拟真实生成误差,有效保留整体结构。
  • 基于关键点匹配评估细节保真度,超越传统语义相似度衡量。

我们提出 FlowFixer,一种针对主体驱动生成(SDG)的精炼框架,用于恢复因主体尺度和视角变化导致的细节丢失。该方法采用视觉参考的图像到图像直接转换,避免语言提示带来的模糊性。为支持图像到图像训练,我们引入一步去噪方案生成自监督训练数据,能自动去除高频细节而保留全局结构,有效模拟真实世界中的SDG误差。我们进一步提出基于关键点匹配的评估指标,以更准确衡量细节保真度,超越通常由CLIP或DINO衡量的语义相似性。实验表明,FlowFixer在定性和定量评估中均优于现有最先进方法,为高保真主体驱动生成设立了新基准。

原文摘要 · Abstract (English)

We present FlowFixer, a refinement framework for subject-driven generation (SDG) that restores fine details lost during generation caused by changes in scale and perspective of a subject. FlowFixer proposes direct image-to-image translation from visual references, avoiding ambiguities in language prompts. To enable image-to-image training, we introduce a one-step denoising scheme to generate self-supervised training data, which automatically removes high-frequency details while preserving global structure, effectively simulating real-world SDG errors. We further propose a keypoint matching-based metric to properly assess fidelity in details beyond semantic similarities usually measured by CLIP or DINO. Experimental results demonstrate that FlowFixer outperforms state-of-the-art SDG methods in both qualitative and quantitative evaluations, setting a new benchmark for high-fidelity subject-driven generation.

图像生成细节修复主体驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。