arXiv:2606.13676cs.CV2026-06被引 2

用简单方法让图像模型同时生成图像和深度图,效果超越现有模型。

Modality Forcing for Scalable Spatial Generation

论文配图:Modality Forcing for Scalable Spatial Generation
图 1 · 摘自论文原文
  • 给不同模态分配独立噪声水平,实现图像与深度联合生成
  • 在稀疏真实深度数据上训练,大模型生成深度更准
  • 适合想低成本构建多模态生成系统的研究者

文本到图像(T2I)模型蕴含丰富的空间先验。生成逼真、复杂的场景需理解几何关系,包括透视和相对尺度。已有工作尝试利用T2I模型进行深度预测,但依赖密集深度数据且流程复杂。本文提出Modality Forcing,一种简单的后训练方案,仅用一个在稀疏深度数据上训练的DiT模型,即可实现图像与深度的任意条件或联合生成。通过为每种模态分配独立噪声水平,并使用分模态解码器,可在真实世界稀疏深度数据上训练,获得强泛化能力的深度预测。进一步实验表明,Modality Forcing继承了T2I预训练的可扩展性:从头训练一系列参数量从370M到3.3B的T2I模型,发现更大模型在更多图像数据上训练时,深度预测更准确。最强模型相比现有联合生成模型,绝对相对误差(AbsRel)降低57%,性能媲美顶尖单目深度估计器。结果有力证明图像生成是空间感知的可扩展预训练目标。

原文摘要 · Abstract (English)

Text-to-image (T2I) models contain rich spatial priors. Synthesizing photorealistic, cluttered scenes requires an understanding of geometry, including perspective and relative scale. Prior works adapt T2I models to leverage this prior for depth prediction, but they require dense depth data and involve complex recipes. We propose Modality Forcing, a simple, scalable post-training recipe for joint image-depth generation using a single DiT trained on sparse depth data. Modality Forcing enables conditional and joint generation of image and depth in any permutation by assigning separate noise levels per modality. Per-modality decoders let us train on sparse, real-world depth and achieve strong, generalizable depth prediction. We further show that Modality Forcing inherits the scalability of T2I pre-training: by training a set of T2I models from scratch (370M to 3.3B parameters), we find that larger models trained on more image data produce more accurate depth. Our strongest model is competitive with state-of-the-art monocular depth estimators and reduces AbsRel by 57% relative to existing joint image-depth generative models. These results provide strong evidence that image generation is a scalable pre-training objective for spatial perception. https://modality-forcing.github.io/

图像生成深度估计多模态可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。