通过动态选择扩散过程中的最佳时间步,提升零样本分割精度。
Unlocking Diffusion Hierarchies: Adaptive Timestep Selection for Zero-Shot Segmentation

- 融合高分辨率注意力图与编码器特征,生成精细像素表征。
- 发现扩散过程存在从局部到整体的语义层次演进规律。
- 按像素自适应选择最优时间步,显著优于现有方法。
零样本分割近年来通过利用大规模文本到图像扩散模型(如Stable Diffusion)中的丰富视觉先验取得了显著进展。然而,现有基于扩散的方法常受限于空间分辨率与上下文信息之间的权衡,且依赖单一静态时间步进行特征提取。为此,本文提出两项关键改进:首先,通过上下文相似性图融合高分辨率注意力图与丰富的U-Net编码器特征,提供细粒度且鲁棒的逐像素表示;其次,我们识别出各类扩散模型在去噪过程中存在一种涌现的层次化语义演进:早期时间步呈现部件级抽象,后期则转向物体级抽象。基于此,我们设计了一种机制,实现对每个像素自适应选择最优时间步。大量实验表明,该方法持续超越现有零样本分割基线,验证了结合上下文特征与动态层级时间步选择的有效性。
原文摘要 · Abstract (English)
Zero-shot segmentation has recently shown notable improvement by leveraging the rich visual priors in large-scale text-to-image diffusion models, such as Stable Diffusion. However, current diffusion-based methods often face limitations due to the trade-off between spatial resolution and contextual information, as well as their reliance on a single static timestep for feature extraction. To overcome these challenges, our work introduces two key advancements. First, our Contextual Similarity Maps fuse high-resolution attention maps with rich U-Net encoder features, providing both fine-grained and robust per-pixel representations. Second, we identify an emergent hierarchical semantic progression within the denoising process of various diffusion models: representations transition from part-level abstractions at earlier timesteps to object-level abstractions at later stages. Leveraging this insight, we introduce a mechanism to adaptively select the optimal timestep for each pixel. Extensive experiments demonstrate that our method consistently outperforms existing zero-shot segmentation baselines, validating the efficacy of combining contextual features with dynamic, hierarchical timestep selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。