用扩散模型生成合成数据,结合令牌解耦提升3D新视角合成效果。
Scaling Transformer-Based Novel View Synthesis Models with Token Disentanglement and Synthetic Data
- 在Transformer中引入令牌解耦机制,分离特征提升表达能力。
- 结合扩散模型合成数据,使模型在未见场景中表现更优。
- 大幅降低计算成本,适合大规模3D内容生成应用。
基于大容量Transformer的模型在稀疏输入视角下实现了可泛化的新型视图合成,无需测试时优化即可生成新视角。然而,这些模型受限于公开场景数据集的多样性,导致大多数真实世界(in-the-wild)场景处于分布外。为此,我们引入由扩散模型生成的合成数据,提升模型在未知域上的泛化能力。尽管合成数据具有可扩展性,但生成过程中引入的伪影成为影响重建质量的关键瓶颈。为此,我们在Transformer架构中提出一种令牌解耦过程,增强特征分离,实现更有效的学习。该改进不仅提升了重建质量,还支持与合成数据的大规模训练。结果表明,该方法在同数据集和跨数据集评估中均优于现有模型,在多个基准上达到最新水平,同时显著降低计算开销。
原文摘要 · Abstract (English)
Large transformer-based models have made significant progress in generalizable novel view synthesis (NVS) from sparse input views, generating novel viewpoints without the need for test-time optimization. However, these models are constrained by the limited diversity of publicly available scene datasets, making most real-world (in-the-wild) scenes out-of-distribution. To overcome this, we incorporate synthetic training data generated from diffusion models, which improves generalization across unseen domains. While synthetic data offers scalability, we identify artifacts introduced during data generation as a key bottleneck affecting reconstruction quality. To address this, we propose a token disentanglement process within the transformer architecture, enhancing feature separation and ensuring more effective learning. This refinement not only improves reconstruction quality over standard transformers but also enables scalable training with synthetic data. As a result, our method outperforms existing models on both in-dataset and cross-dataset evaluations, achieving state-of-the-art results across multiple benchmarks while significantly reducing computational costs. Project page: https://scaling3dnvs.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。