arXiv:2604.01761cs.CV2026-04

用自监督特征控制视频生成,实现风格迁移和3D转视频。

Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion

论文配图:Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion
图 1 · 摘自论文原文
  • 用DINOv3特征解耦外观与语义,实现可控视频生成
  • 低分辨率空间信息可用高维特征补偿,提升控制精度
  • 适合需要精细外观控制的视频生成与3D内容转换任务

视频扩散模型在内容生成、新视角合成和世界模拟中表现优异。许多生成与迁移应用依赖于感知、几何或简单语义信号进行条件控制,本质上将其作为生成渲染器使用。与此同时,通过大规模自监督学习(如DINOv3)获得的高维特征正成为视觉模型的通用接口。尽管已有研究探索其在特定主体编辑和模型对齐中的应用,但尚未作为密集条件信号用于预训练视频扩散模型。这些自监督特征包含大量关于场景风格、光照和语义的纠缠信息,虽利于重建,却限制了生成能力。本文提出一种轻量级控制架构与训练策略,可解耦外观与其他需保留的特征,实现风格化与再照明等外观变化的稳健控制。此外,我们证明低空间分辨率可通过更高维度特征补偿,从而提升从显式空间表示生成视频时的可控性。

原文摘要 · Abstract (English)

Video diffusion models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically through perceptual, geometric, or simple semantic signals, fundamentally using them as generative renderers. At the same time, high-dimensional features obtained from large-scale self-supervised learning on images or point clouds are increasingly used as a general-purpose interface for vision models. The connection between the two has been explored for subject specific editing, aligning and training video diffusion models, but not in the role of a dense conditioning signal for pretrained video diffusion models. Features obtained through self-supervised learning like DINOv3, contain a lot of entangled information about style, lighting and semantics of the scene. This makes them great at reconstruction tasks but limits their generative capabilities. In this paper, we show how we can use the features for tasks such as video domain transfer and video-from-3D generation. We introduce a lightweight control architecture and training strategy that decouples appearance from other features that we wish to preserve, enabling robust control for appearance changes such as stylization and relighting. Furthermore, we show that low spatial resolution can be compensated by higher feature dimensionality, improving controllability in generative rendering from explicit spatial representations.

视频生成扩散模型特征控制3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。