arXiv:2512.03045cs.CV2025-12被引 6

通过监督注意力图对齐,提升多视角生成的几何一致性。

CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models

  • 用几何对应关系直接监督注意力层,增强模型对齐能力。
  • 仅监督一层注意力即可使收敛速度翻倍,生成质量更优。
  • 方法通用,适用于任意多视角扩散模型,无需修改结构。

多视角扩散模型在新视角合成中表现强劲,但其保持视图一致性的内在机制尚不明确。本文首次验证,训练过程中注意力图会自发学习到跨参考与目标视角的几何对应区域。然而,在大视角变化下,这种对应信号仍不完整。基于此,我们提出CAMEO——一种简单有效的训练技术,通过几何对应关系直接监督注意力图,显著提升多视角扩散模型的训练效率与生成质量。关键发现:仅需监督单一注意力层,即可引导模型学习精确对应关系,从而保留参考图像的几何结构,加速收敛,并提升新视角合成性能。实验表明,CAMEO可将收敛所需迭代次数减少一半,且在相同迭代数下表现更优。此外,该方法具有模型无关性,可无缝应用于任意多视角扩散模型。

原文摘要 · Abstract (English)

Multi-view diffusion models have recently emerged as a powerful paradigm for novel view synthesis, yet the underlying mechanism that enables their view-consistency remains unclear. In this work, we first verify that the attention maps of these models acquire geometric correspondence throughout training, attending to the geometrically corresponding regions across reference and target views for view-consistent generation. However, this correspondence signal remains incomplete, with its accuracy degrading under large viewpoint changes. Building on these findings, we introduce CAMEO, a simple yet effective training technique that directly supervises attention maps using geometric correspondence to enhance both the training efficiency and generation quality of multi-view diffusion models. Notably, supervising a single attention layer is sufficient to guide the model toward learning precise correspondences, thereby preserving the geometry and structure of reference images, accelerating convergence, and improving novel view synthesis performance. CAMEO reduces the number of training iterations required for convergence by half while achieving superior performance at the same iteration counts. We further demonstrate that CAMEO is model-agnostic and can be applied to any multi-view diffusion model.

扩散模型多视角生成注意力监督几何对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。