arXiv:2511.00293cs.CV2025-11

用3D先验引导上下文学习,实现少样本多视角身份一致生成

MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning

  • 基于3D先验的上下文学习架构,激活模型多视角生成能力
  • 仅用100张多视角样本,即在多视角一致性上超越主流方法
  • 适合需要少样本、多视角身份定制的图像生成应用

近期个性化生成模型在跨场景生成同一人物身份一致图像方面表现卓越,但多数方法缺乏显式的视角控制,难以保证生成身份的多视角一致性。为此,我们提出MagicView,一种轻量级适配框架,通过3D先验引导的上下文学习,赋予现有生成模型多视角生成能力。尽管先前研究显示上下文学习可保持网格采样下的身份一致性,其在多视角场景下的有效性尚未被探索。基于此,我们深入分析了多视角上下文学习能力,并设计了一种利用3D先验激活该能力的条件架构。同时,获取稳健的多视角能力通常依赖大规模多维数据集,导致在小样本条件下文本可控性易退化。为此,我们引入一种新型语义对应对齐损失,有效维持语义一致性与多视角一致性。大量实验表明,MagicView在多视角一致性、文本对齐、身份相似性和视觉质量上显著优于近期基线,在仅使用100个多视角训练样本下取得优异表现。

原文摘要 · Abstract (English)

Recent advances in personalized generative models have demonstrated impressive capabilities in producing identity-consistent images of the same individual across diverse scenes. However, most existing methods lack explicit viewpoint control and fail to ensure multi-view consistency of generated identities. To address this limitation, we present MagicView, a lightweight adaptation framework that equips existing generative models with multi-view generation capability through 3D priors-guided in-context learning. While prior studies have shown that in-context learning preserves identity consistency across grid samples, its effectiveness in multi-view settings remains unexplored. Building upon this insight, we conduct an in-depth analysis of the multi-view in-context learning ability, and design a conditioning architecture that leverages 3D priors to activate this capability for multi-view consistent identity customization. On the other hand, acquiring robust multi-view capability typically requires large-scale multi-dimensional datasets, which makes incorporating multi-view contextual learning under limited data regimes prone to textual controllability degradation. To address this issue, we introduce a novel Semantic Correspondence Alignment loss, which effectively preserves semantic alignment while maintaining multi-view consistency. Extensive experiments demonstrate that MagicView substantially outperforms recent baselines in multi-view consistency, text alignment, identity similarity, and visual quality, achieving strong results with only 100 multi-view training samples.

多视角生成身份一致小样本学习3D先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。