arXiv:2606.29198cs.CV2026-06中稿 · ECCV

通过动态轨迹初始化提升生成式人脸视频超分的保真度与效率

DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution

论文配图:DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution
图 1 · 摘自论文原文
  • 将生成式超分重构成输入驱动的定向修复,利用条件增强机制提升质量
  • 仅需微调模型,即在多个基准上达到当前最优性能,且推理成本更低
  • 提出判别性引导机制,实现信号噪声比对齐,有效缓解感知-失真权衡问题

作为最具感知力的人脸视频超分辨率方法,现有生成式人脸视频超分辨率(GFVSR)主要依赖预训练扩散模型的生成先验。然而,作为完整生成过程,其存在固定采样路径和高昂推理开销的问题,且缺乏大规模辅助训练时尤为明显。此外,过度追求通用感知指标常导致保真度下降。为此,本文提出动态轨迹初始化(DTI)范式,将GFVSR重构为输入驱动的定向修复任务。通过为预训练DiT主干设计新颖的增强-注入条件机制,显著提升了模型保真度而不损失感知质量。为动态设定起始采样点,提出基于目标信噪比(SNR)对齐训练的判别性引导(DG)。仅需少量模型适配与微调,本方法在多个指标和基准上实现当前最优综合表现。进一步分析了实际综合质量与常见指标的关系,揭示感知-失真权衡现象,并验证LPIPS为最可信评价指标。

原文摘要 · Abstract (English)

As the most perceptually powerful Face Video Super-Resolution (FVSR) method, existing works in Generative FVSR (GFVSR) mainly exploit the generative prior of pretrained diffusion models. However, viewed as full generation, they suffer from fixed sampling and expensive inference costs if without large-scale auxiliary training. Furthermore, an excessive pursuit of generic perceptual metrics often results in low fidelity. To address these issues, we present Dynamic Trajectory Initialization (DTI) paradigm for GFVSR, which reformulates GFVSR as an input-driven directional restoration. With a novel enhancement-and-injection conditioning mechanism for pretrained DiT backbone, fidelity of our model has been significantly improved without compromising perceptual quality. To dynamically set the starting sampling point, we propose a Discriminative Guide (DG) trained via objective Signal-to-Noise Ratio (SNR) alignment. With only minor model adaptation and fine-tuning, our method achieves a SOTA overall performance across diverse metrics and benchmarks. An analysis of relationship between actual comprehensive quality and common metrics is also conducted, which demonstrates the perception-distortion trade-off and that the LPIPS is the most convincing metric in our case.

视频超分扩散模型生成对抗轨迹初始化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。