用物理光学模拟解决视频戴眼镜去除中的形变与失真问题。
Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal

- 通过物理光学仿真生成配对数据,增强模型对镜片折射的建模能力。
- 在FFHQ数据集上达到27.68 FPS推理速度,保持高保真与结构准确。
- 适合需要高一致性视频人脸编辑的研究者与应用开发者。
从视频中高保真地去除眼镜是面部属性编辑的重大挑战,因镜片引起的复杂折射畸变和视角依赖的镜面反射常掩盖真实面部几何。尽管大规模生成先验在静态图像修复中表现良好,但往往缺乏维持身份、表情和姿态的结构约束,导致静态与动态序列中出现明显的“身份漂移”。本文提出一种新型迁移框架:首先从商用生成模型(Nano Banana, Gemini 3 Pro Image)中提取高保真合成人脸图像,经三阶段结构滤波过程正则化以保留身份、表情与姿态,再在训练中引入基于物理的镜片光学仿真,生成多样化的成对数据。该过程将Nano Banana的多视角逼真知识迁移至专用恢复架构JFSnet(联合特征-空间网络)。JFSnet融合DINOv2语义特征与卷积解码器进行空间重建,利用平移等变性约束提升时序一致性与高频细节保留。在精选的Flickr-Faces-HQ(FFHQ)子集(12,163张图像)上的评估表明,本方法实现高保真与结构准确性,同时保持27.68 FPS的推理速度。在CelebV-Text视频序列的感知评测中,结果在眼区一致性、时序稳定性与整体修复质量上均优于扩散模型与GAN基线。
原文摘要 · Abstract (English)
High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections. While large-scale generative priors have shown promise in eye-glasses removal via static image inpainting, they often lack the structural constraints necessary to maintain identity, expression, and pose, leading to visible "identity drift" in both static images and dynamic sequences. In this paper, we propose a novel transfer framework that addresses the stochastic nature of generative priors. Our pipeline first extracts high-fidelity synthetic face images from a commercial-grade generative model (Nano Banana, Gemini 3 Pro Image), regularizes them via a three-stage structural filtering process to preserve identity, expression, and pose, and finally applies physically-based simulation of lens optics during training to provide diverse, paired data. This process transfers Nano Banana's photo-realistic, multi-view knowledge into a specialized restoration architecture, JFSnet (Joint Feature-Spatial network). JFSnet integrates DINOv2-based semantic features with a convolutional decoder for spatial reconstruction, leveraging translation equivariance constraints to improve temporal consistency and high-frequency detail preservation. Evaluations on the curated Flickr-Faces-HQ (FFHQ) subset (12,163 images) show that our approach achieves high fidelity and structural accuracy, while maintaining inference speed of 27.68 FPS. In perceptual studies on CelebV-Text video sequences, our results are consistently preferred over diffusion and GAN-based baselines for ocular consistency, temporal stability, and overall restoration quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。