无需标注数据,一键实现视频光影色彩自由编辑
VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning

- 利用自监督学习分离视频结构与光影色彩,自动重组生成
- 零样本适配多种编辑任务,视觉质量高且计算开销低
- 适合视频制作、影视后期等需要快速调色的场景
视频色彩光照编辑旨在调整光照与色彩同时保持结构和时序一致性,仍是重大挑战。现有方法多依赖昂贵的有监督训练及合成成对数据。本文提出VibeFlow,一种新颖的自监督框架,挖掘预训练视频生成模型的内在物理理解能力。不从头学习色彩光照变化,而是引入解耦数据扰动流程,强制模型自适应地融合源视频结构与参考图像的色彩-光照信息,实现鲁棒的无监督解耦。此外,为修正流模型固有的离散化误差,引入残差速度场与结构失真一致性正则化,确保严格结构保真与时序连贯性。该框架无需昂贵训练资源,可零样本泛化至多种应用,包括视频重光照、重着色、低光增强、昼夜转换及对象级颜色编辑。大量实验表明,VibeFlow在显著降低计算开销的同时达到优异视觉效果。
原文摘要 · Abstract (English)
Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge. Existing methods typically rely on expensive supervised training with synthetic paired data. This paper proposes VibeFlow, a novel self-supervised framework that unleashes the intrinsic physical understanding of pre-trained video generation models. Instead of learning color and light transitions from scratch, we introduce a disentangled data perturbation pipeline that enforces the model to adaptively recombine structure from source videos and color-illumination cues from reference images, enabling robust disentanglement in a self-supervised manner. Furthermore, to rectify discretization errors inherent in flow-based models, we introduce Residual Velocity Fields alongside a Structural Distortion Consistency Regularization, ensuring rigorous structural preservation and temporal coherence. Our framework eliminates the need for costly training resources and generalizes in a zero-shot manner to diverse applications, including video relighting, recoloring, low-light enhancement, day-night translation, and object-specific color editing. Extensive experiments demonstrate that VibeFlow achieves impressive visual quality with significantly reduced computational overhead. Our project is publicly available at https://lyf1212.github.io/VibeFlow-webpage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。