用图像编辑模型做密集感知,效果更准更快。
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
- 基于图像编辑扩散模型,实现端到端的结构保持优化。
- 在深度、法线、抠图任务上均达最新最好性能。
- 单步推理速度快,小数据训练也有效,适合实用场景。
近期扩散变压器在视觉合成方面表现出色,但大多数密集感知方法仍依赖为随机生成设计的文本到图像(T2I)生成器。本文重新审视这一范式,发现图像编辑扩散模型天然具备图像到图像的一致性,更适合密集感知任务。提出 Edit2Perceive,一个统一的扩散框架,将编辑模型用于深度估计、法线预测和抠图任务。基于 FLUX.1 Kontext 架构,采用全参数微调与像素空间一致性损失,确保中间去噪过程中的结构保持。此外,单步确定性推理实现最高加速比,且在相对较小的数据集上训练即可达到优异效果。大量实验表明,该方法在三项任务上均取得全面领先,揭示了面向编辑的扩散变压器在几何感知方面的强大潜力。
原文摘要 · Abstract (English)
Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm and show that image editing diffusion models are inherently image-to-image consistent, providing a more suitable foundation for dense perception task. We introduce Edit2Perceive, a unified diffusion framework that adapts editing models for depth, normal, and matting. Built upon the FLUX.1 Kontext architecture, our approach employs full-parameter fine-tuning and a pixel-space consistency loss to enforce structure-preserving refinement across intermediate denoising states. Moreover, our single-step deterministic inference yields up to faster runtime while training on relatively small datasets. Extensive experiments demonstrate comprehensive state-of-the-art results across all three tasks, revealing the strong potential of editing-oriented diffusion transformers for geometry-aware perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。