arXiv:2508.15774cs.CV2025-08被引 11

无需微调即可生成8K高清图像,4K视频仅需少量微调。

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

  • 提出CineScale推理范式,解决高分辨率生成中的高频误差累积问题。
  • 实现8K图像零微调生成,4K视频仅需最小LoRA微调。
  • 支持图像到视频、视频到视频等多种生成任务,适配主流开源框架。

视觉扩散模型虽取得显著进展,但通常受限于低分辨率训练数据与计算资源,难以生成高保真图像或视频。近期研究尝试通过无微调策略挖掘预训练模型的高分辨率生成潜力,但易产生重复模式等低质量内容。根本原因在于超出训练分辨率时高频信息必然增加,导致误差累积引发重复图案。本文提出CineScale,一种新型推理范式,针对两类视频生成架构设计专用变体。不同于仅限于高分辨率文本到图像(T2I)和文本到视频(T2V)的基线方法,CineScale拓展至高分辨率图像到视频(I2V)和视频到视频(V2V)合成,基于当前最先进的开源视频生成框架构建。大量实验验证了该范式在提升图像与视频模型高分辨率生成能力方面的优越性。显著成果包括:无需任何微调即可实现8K图像生成,且仅需极小量LoRA微调即可达成4K视频生成。生成视频样本可访问官网:https://eyeline-labs.github.io/CineScale/。

原文摘要 · Abstract (English)

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity images or videos at higher resolutions. Recent efforts have explored tuning-free strategies to exhibit the untapped potential higher-resolution visual generation of pre-trained models. However, these methods are still prone to producing low-quality visual content with repetitive patterns. The key obstacle lies in the inevitable increase in high-frequency information when the model generates visual content exceeding its training resolution, leading to undesirable repetitive patterns deriving from the accumulated errors. In this work, we propose CineScale, a novel inference paradigm to enable higher-resolution visual generation. To tackle the various issues introduced by the two types of video generation architectures, we propose dedicated variants tailored to each. Unlike existing baseline methods that are confined to high-resolution T2I and T2V generation, CineScale broadens the scope by enabling high-resolution I2V and V2V synthesis, built atop state-of-the-art open-source video generation frameworks. Extensive experiments validate the superiority of our paradigm in extending the capabilities of higher-resolution visual generation for both image and video models. Remarkably, our approach enables 8k image generation without any fine-tuning, and achieves 4k video generation with only minimal LoRA fine-tuning. Generated video samples are available at our website: https://eyeline-labs.github.io/CineScale/.

视频生成扩散模型高分辨率无微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。