arXiv:2608.04224cs.CV2026-08被引 1

首个联合音视频修复模型,同步恢复老电影画质与音质。

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

论文配图:OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
图 1 · 摘自论文原文
  • 用统一多模态扩散模型,联合条件生成修复音视频
  • 在200段真实老片上超越所有方法,实现自然着色与音画同步
  • 适合影视修复、文化遗产数字化领域研究者

历史影片常同时存在画质模糊、噪声、闪烁和音频嘶声、削波、断续等退化问题,现有方法独立修复音视频,导致质量差距与跨模态不一致。本文提出OmniVR,首个联合音视频生成式修复模型。基于220亿参数的音视频生成主干网络,将修复建模为统一多模态DiT中的条件生成:低质量音视频作为潜在条件,结合固定修复提示,联合去噪以恢复视觉结构、时间运动与声学细节。三个关键设计实现该目标:(1) 基于互联网收集数据模拟真实老片退化特征的联合音视频退化流程;(2) 保持架构的文本到音视频(T2AV)至音视频到音视频(AV2AV)过渡,配合提示渐进调节,最大限度保留生成先验;(3) 首帧图像到视频(I2V)锚定结合损失重加权与波形监督,提升长视频外推与音频保真度。此外,提出OmniVRBench,首个评估音视频修复在视觉质量、音频质量、时间一致性与音画同步性方面的基准,覆盖200段真实历史片段。OmniVR在所有六项视觉指标上均优于此前方法,音频质量最佳,并首次实现自然着色。代码与权重将公开发布。

原文摘要 · Abstract (English)

Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. We present OmniVR, the first joint audio-video generative restoration model. Built upon a 22B-parameter audio-video generation backbone, OmniVR formulates restoration as conditional generation within a unified multimodal DiT: the low-quality video and audio are encoded as latent conditions, combined with a fixed restoration prompt, and jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. Three key designs enable this adaptation: (1) a joint audio-video degradation pipeline that simulates real old-film characteristics from Internet-collected data; (2) an architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally retains the generative prior; and (3) first-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity. We also propose OmniVRBench, the first benchmark that evaluates audio-video restoration across visual quality, audio quality, temporal consistency, and audio-visual synchrony on 200 real historical clips. OmniVR surpasses all prior methods on all six visual metrics, achieves the best audio quality, and produces natural colorization---the first method to jointly address all three aspects. Code and weights will be publicly released. Project Page: https://xin1u.github.io/OminiVR_PAGE/

音视频修复老片复原多模态生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。