改进扩散模型生成分割的向量场学习,提升收敛速度与类别区分度。
Rethinking Vector Field Learning for Generative Segmentation
- 通过引入距离感知修正项重塑速度场,增强梯度强度。
- 在多个数据集上显著提升分割质量,接近判别式模型性能。
- 适合关注生成式分割与扩散模型优化的研究者。
将扩散模型用于生成式分割受到越来越多关注。现有方法多集中于架构调整或训练技巧,但对连续流匹配目标与离散感知任务之间的内在不匹配理解有限。本文从向量场学习视角重新审视扩散分割,发现常用流匹配目标存在梯度消失与轨迹穿越两大问题,导致收敛慢、类别分离差。为此,提出一种基于原理的向量场重塑策略,通过加入独立的距离感知修正项,引入吸引与排斥作用,在保持原始扩散训练框架的同时,增强中心点附近的梯度幅度。此外,设计了一种受Kronecker序列启发的计算高效的准随机类别编码方案,可无缝集成至端到端像素神经场框架中实现像素级语义对齐。大量实验表明,该方法在多个基准上显著优于基线流匹配方法,大幅缩小生成式分割与强判别式模型之间的性能差距。
原文摘要 · Abstract (English)
Taming diffusion models for generative segmentation has attracted increasing attention. While existing approaches primarily focus on architectural tweaks or training heuristics, there remains a limited understanding of the intrinsic mismatch between continuous flow matching objectives and discrete perception tasks. In this work, we revisit diffusion segmentation from the perspective of vector field learning. We identify two key limitations of the commonly used flow matching objective: gradient vanishing and trajectory traversing, which result in slow convergence and poor class separation. To tackle these issues, we propose a principled vector field reshaping strategy that augments the learned velocity field with a detached distance-aware correction term. This correction introduces both attractive and repulsive interactions, enhancing gradient magnitudes near centroids while preserving the original diffusion training framework. Furthermore, we design a computationally efficient, quasi-random category encoding scheme inspired by Kronecker sequences, which integrates seamlessly with an end-to-end pixel neural field framework for pixel-level semantic alignment. Extensive experiments consistently demonstrate significant improvements over vanilla flow matching approaches, substantially narrowing the performance gap between generative segmentation and strong discriminative specialists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。