arXiv:2512.14068cs.CVcs.AI2025-12ACL被引 10

提出首个高效稳定的块级扩散模型,显著提升视觉语言理解的训练效率与性能。

SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding

  • 采用异步块噪声调度、掩码比例归一化与渐进噪声课程,提升训练稳定性。
  • 在21个基准上实现比传统块扩散更快收敛与更高精度,部分超越自回归模型。
  • 适合追求高效视觉语言建模的开发者与研究者使用。

块级离散扩散在并行生成与因果依赖建模间取得良好平衡,是视觉语言建模的有前景框架。然而,高训练成本、慢收敛与不稳定性限制了其实际应用,使其长期落后于强自回归基线。本文提出SDAR-VL,首次系统性将块级离散扩散应用于大规模视觉语言理解(VLU),并设计了一个集成式高效稳定训练框架。该框架包含三项核心技术:(1) 异步块级噪声调度,增强批内监督多样性;(2) 有效掩码比例缩放,实现随机掩码下的无偏损失归一化;(3) 渐进β噪声课程,提升有效掩码覆盖率同时保持扰动多样性。在21个单图、多图及视频基准上的实验表明,SDAR-VL在训练效率、收敛稳定性与任务性能上均优于传统块扩散。在该评测集上,其成为扩散类视觉语言模型的新SOTA,且在匹配设置下达到或超过强自回归基线LLaVA-OneVision与全局扩散基线LLaDA-V,确立块级扩散作为视觉语言理解实用骨干的可行性。

原文摘要 · Abstract (English)

Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been limited by high training cost, slow convergence, and instability, which have so far kept it behind strong autoregressive (AR) baselines. We present \textbf{SDAR-VL}, the first systematic application of block-wise discrete diffusion to large-scale vision-language understanding (VLU), together with an \emph{integrated framework for efficient and stable training}. This framework unifies three components: (1) \textbf{Asynchronous Block-wise Noise Scheduling} to diversify supervision within each batch; (2) \textbf{Effective Mask Ratio Scaling} for unbiased loss normalization under stochastic masking; and (3) a \textbf{Progressive Beta Noise Curriculum} that increases effective mask coverage while preserving corruption diversity. Experiments on 21 single-image, multi-image, and video benchmarks show that SDAR-VL consistently improves \emph{training efficiency}, \emph{convergence stability}, and \emph{task performance} over conventional block diffusion. On this evaluation suite, SDAR-VL sets a new state of the art among diffusion-based vision-language models and, under matched settings, matches or surpasses strong AR baselines such as LLaVA-OneVision as well as the global diffusion baseline LLaDA-V, establishing block-wise diffusion as a practical backbone for VLU.

视觉语言扩散模型块级扩散高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。