arXiv:2608.29233cs.CV2026-08

通过分离验证场景避免记忆偏差,提升单图多视角生成泛化能力。

Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis

论文配图:Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis
图 1 · 摘自论文原文
  • 用不重叠的验证场景防止模型记忆训练数据
  • 在40个场景下实现26个视角的高质量生成
  • 适合小样本生成任务和对泛化性要求高的研究者

我们提出应对ACM Multimedia 2026 Grand Challenge的优胜方案,该挑战要求仅用40个训练场景,从一张RGB图像生成26个目标视角,每视图仅允许一次前向传播。禁止使用显式几何、外部渲染、链式生成、候选选择和后处理。我们发现:共享训练与验证场景会导致模型将记忆误作可迁移的视角控制。为此提出GoM(Generalization over Memorization)框架,结合场景不重叠验证、曝光匹配选择与靶向扩散适配。其合成模型采用4B参数的修正流DiT,通过秩为32的LoRA进行适配,配合优化器重启、晚期检查点平均及VAE解码器调优。超过300次离线实验与24次在线提交表明,在小样本生成中,验证设计与训练轨迹控制的重要性不亚于模型规模。

原文摘要 · Abstract (English)

We present the winning solution to the ACM Multimedia 2026 Grand Challenge on Single-Image Guided Multi-Angle Image Synthesis. It ranks first among 293 registered teams; 56 teams obtained at least one scored submission on the public Phase-A leaderboard. With only 40 training scenes, the challenge requires 26 target views from one RGB model and one forward pass per view; it prohibits explicit geometry, external rendering, chained generation, candidate selection, and post-processing. We identify a critical model-selection failure: shared training and validation scenes make memorization appear as transferable view control. We therefore introduce GoM. Short for Generalization over Memorization, the framework combines scene-disjoint validation, exposure-matched selection, and targeted diffusion adaptation. Its synthesis model adapts a 4B rectified-flow DiT using rank-32 LoRA, optimizer restarts, late-checkpoint averaging, and VAE decoder tuning. More than 300 offline experiments and 24 online submissions show that validation design and training-trajectory control can matter as much as architecture scale in small-data generative modeling.

扩散模型单图生成泛化性小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。