arXiv:2602.10113cs.CV2026-02被引 6

让图像生成视频时保持物体身份一致,视角变化也不变形。

ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation

论文配图:ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
图 1 · 摘自论文原文
  • 用多视角辅助图增强单图信息,融合语义与几何特征
  • 在真实场景下,身份保真度和时间连贯性显著优于现有模型
  • 适用于需要精准物体一致性生成的视频应用

图像到视频生成(I2V)将静态图像转化为遵循文本指令的连贯视频序列,但如何在视角变化下保持细粒度物体身份一致仍是难题。与文本到视频模型不同,现有I2V方法常出现外观漂移和几何失真,我们归因于单视图2D观测稀疏及跨模态对齐弱。为此,我们从数据与模型双角度出发:首先构建了大规模物体中心数据集ConsIDVid,采用可扩展流程生成高质量、时间对齐视频,并建立ConsIDVid-Bench,提出针对多视角一致性的新评估框架,可检测细微几何与外观偏差。进一步提出ConsID-Gen,一种视图辅助的I2V生成框架,通过在首帧附加无姿态辅助视图,利用双流视觉-几何编码器融合语义与结构线索,并通过文本-视觉连接器统一条件输入至扩散变换器主干网络。在ConsIDVid-Bench上的实验表明,ConsID-Gen在多个指标上持续领先,最佳性能超越Wan2.1和HunyuanVideo等先进视频生成模型,在复杂真实场景中实现更优的身份保真度与时间连贯性。代码与数据集将于https://myangwu.github.io/ConsID-Gen发布。

原文摘要 · Abstract (English)

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike text-to-video models, existing I2V pipelines often suffer from appearance drift and geometric distortion, artifacts we attribute to the sparsity of single-view 2D observations and weak cross-modal alignment. Here we address this problem from both data and model perspectives. First, we curate ConsIDVid, a large-scale object-centric dataset built with a scalable pipeline for high-quality, temporally aligned videos, and establish ConsIDVid-Bench, where we present a novel benchmarking and evaluation framework for multi-view consistency using metrics sensitive to subtle geometric and appearance deviations. We further propose ConsID-Gen, a view-assisted I2V generation framework that augments the first frame with unposed auxiliary views and fuses semantic and structural cues via a dual-stream visual-geometric encoder as well as a text-visual connector, yielding unified conditioning for a Diffusion Transformer backbone. Experiments across ConsIDVid-Bench demonstrate that ConsID-Gen consistently outperforms in multiple metrics, with the best overall performance surpassing leading video generation models like Wan2.1 and HunyuanVideo, delivering superior identity fidelity and temporal coherence under challenging real-world scenarios. We will release our model and dataset at https://myangwu.github.io/ConsID-Gen.

图像转视频身份保持多视角生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。