arXiv:2603.06507cs.CV2026-03被引 15

自监督流匹配让生成模型同时学会语义表示,提升多模态合成质量。

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

  • 通过异步时间调度引入跨标记噪声差异,强制模型推断缺失信息。
  • 在图像、视频、音频上实现更优生成效果,且符合预期缩放规律。
  • 无需外部模型,统一框架支持多模态训练,适合大规模生成任务。

强大的语义表征能提升扩散模型与流模型的收敛速度与生成质量。现有方法大多依赖外部模型,需额外训练,目标不一致,且表现出意外的缩放行为。我们指出,这种依赖源于训练目标——去噪任务缺乏学习语义表征的动力。为此提出 Self-Flow:一种将表征学习融入生成框架的自监督流匹配范式。核心机制为双时间步调度,在不同标记上施加异构噪声水平,制造信息不对称,迫使模型从受损输入中推断缺失内容,从而在无外部监督下同步学习强表征与生成能力。该方法具有跨模态泛化性,支持多模态联合训练,并遵循预期缩放规律,在图像、视频、音频生成上均取得更优表现。

原文摘要 · Abstract (English)

Strong semantic representations improve the convergence and generation quality of diffusion and flow models. Existing approaches largely rely on external models, which require separate training, operate on misaligned objectives, and exhibit unexpected scaling behavior. We argue that this dependence arises from the model's training objective, which poses a denoising task with little incentive to learn semantic representations. We introduce Self-Flow: a self-supervised flow matching paradigm that integrates representation learning within the generative framework. Our key mechanism, Dual-Timestep Scheduling, applies heterogeneous noise levels across tokens, creating an information asymmetry that forces the model to infer missing information from corrupted inputs. This drives learning strong representations alongside generative capabilities without external supervision. Our method generalizes across modalities and enables multi-modal training while following expected scaling laws, achieving superior image, video, and audio generation.

生成模型自监督多模态流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。