arXiv:2603.10785cs.CV2026-03

通过几何视角优化文本到图像生成的训练效率与质量

The Quadratic Geometry of Flow Matching: Semantic Granularity Alignment for Text-to-Image Synthesis

  • 揭示流匹配中损失函数的二次型结构与动态核机制
  • 提出语义粒度对齐方法,加速收敛并提升图像结构完整性
  • 适合关注生成模型训练优化与跨模态生成的研究者

本文分析生成式微调的优化动态,发现流匹配框架下标准MSE目标可表述为由动态演化神经正切核(NTK)控制的二次型。这一几何视角揭示了隐含的数据交互矩阵:对角项表示独立样本学习,非对角项编码异质特征间的残差相关性。尽管标准训练隐式优化这些交叉干扰,但缺乏显式控制;且主流数据同质性假设可能限制模型有效容量。受此启发,我们提出语义粒度对齐(SGA),以文本到图像合成作为验证场景。SGA通过干预向量残差场缓解梯度冲突。在DiT与U-Net架构上的评估表明,SGA显著提升效率-质量权衡,加快收敛速度并增强图像结构保真度。

原文摘要 · Abstract (English)

In this work, we analyze the optimization dynamics of generative fine-tuning. We observe that under the Flow Matching framework, the standard MSE objective can be formulated as a Quadratic Form governed by a dynamically evolving Neural Tangent Kernel (NTK). This geometric perspective reveals a latent Data Interaction Matrix, where diagonal terms represent independent sample learning and off-diagonal terms encode residual correlation between heterogeneous features. Although standard training implicitly optimizes these cross-term interferences, it does so without explicit control; moreover, the prevailing data-homogeneity assumption may constrain the model's effective capacity. Motivated by this insight, we propose Semantic Granularity Alignment (SGA), using Text-to-Image synthesis as a testbed. SGA engineers targeted interventions in the vector residual field to mitigate gradient conflicts. Evaluations across DiT and U-Net architectures confirm that SGA advances the efficiency-quality trade-off by accelerating convergence and improving structural integrity.

文本生成图像生成流匹配优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。