让视频扩散模型生成更真实的镜面反射,解决内容错乱问题。
MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

- 用语义关系蒸馏建模镜中该反射什么内容
- 用几何变换对齐确保反射空间布局正确
- 专为镜面反射设计,适合视频修复与生成任务
视频扩散模型(VDMs)在高质量视频合成方面取得进展,但生成镜面反射仍具挑战,因镜中内容必须与周围场景保持一致。现有VDMs未专门建模场景与镜面的关系,常导致反射内容错误或空间不一致。我们发现镜面反射生成包含两个互补挑战:确定应反射哪些场景内容,以及如何安排反射内容的空间位置。为此提出MirrorWorld,一种反射感知的视频修补框架,在生成过程中建模场景-镜面关系。具体引入语义关系蒸馏(SRD),将冻结视觉基础模型中的关系信息迁移到生成过程,增强可见场景与镜面区域间的语义关联;进一步提出几何变换对齐(GTA),学习一个变换以指导反射内容的空间排列。两者协同作用:SRD决定反映什么,GTA决定如何摆放。为推动研究,我们将四个现有视频镜像数据集重构为统一的反射重建基准。实验表明,MirrorWorld在反射重建质量上优于代表性图像级反射生成方法和强视频修补基线。
原文摘要 · Abstract (English)
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。