用视频扩散模型从模糊图生成焦深序列,实现拍摄后交互式调焦。
Learning to Refocus with Video Diffusion Models
- 基于视频扩散模型生成焦深序列,从单张模糊图像恢复多焦点图像
- 在真实手机场景下表现更优,视觉质量与鲁棒性显著提升
- 适合摄影爱好者和图像编辑者,推动手机摄影后期调焦发展
聚焦是摄影的核心,但自动对焦系统常无法捕捉预期主体,用户也常希望拍摄后调整焦点。我们提出一种基于视频扩散模型的新型后处理调焦方法,仅需一张模糊图像即可生成感知真实的焦深序列(以视频形式呈现),支持交互式调焦,并拓展下游应用。我们发布了一个大规模焦深数据集,覆盖多种真实手机拍摄条件,以支持本工作及未来研究。该方法在复杂场景中持续优于现有技术,在感知质量和鲁棒性上均表现突出,为日常摄影中的焦点编辑能力开辟新路径。代码与数据已公开于 www.learn2refocus.github.io。
原文摘要 · Abstract (English)
Focus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video diffusion models. From a single defocused image, our approach generates a perceptually accurate focal stack, represented as a video sequence, enabling interactive refocusing and unlocking a range of downstream applications. We release a large-scale focal stack dataset acquired under diverse real-world smartphone conditions to support this work and future research. Our method consistently outperforms existing approaches in both perceptual quality and robustness across challenging scenarios, paving the way for more advanced focus-editing capabilities in everyday photography. Code and data are available at www.learn2refocus.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。