arXiv:2605.19398cs.CVcs.AI2026-05被引 2

通过平衡参考帧注意力提升图像生成视频的动态表现

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

论文配图:Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
图 1 · 摘自论文原文
  • 发现非参考帧过度关注参考帧关键信息导致运动受限
  • 提出无训练、免修改模型的DyMoS方法,可调控制运动强度
  • 无需改动原模型,兼顾画面质量与参考图保真度

图像到视频(I2V)模型生成的视频常显得过于静态,相较于文本到视频模型。以往方法通过弱化或修改图像条件信号缓解此问题,但往往需要额外训练或牺牲对参考图像的保真度。本文识别出参考帧主导性是抑制运动的关键机制:I2V模型中非参考帧在自注意力过程中过度聚焦于参考帧的关键令牌,导致参考信息跨时间过量传播,抑制了帧间动态。基于此发现,提出无需训练、模型无关的DyMoS(Dynamic Motion Slider)方法,在初始去噪步骤中重新平衡生成帧到参考帧的注意力路径。DyMoS不改变输入图像和模型权重,仅引入一个标量参数实现运动强度的连续调控。在多个主流I2V骨干网络上实验表明,DyMoS持续提升运动动态性,同时保持视觉质量和参考图像一致性。

原文摘要 · Abstract (English)

Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fidelity to the reference image. In this work, we identify reference-frame dominance as a key mechanism behind motion suppression. We observe that non-reference frames in I2V models allocate excessive self-attention to reference-frame key tokens, causing reference information to be over-propagated across time and suppressing inter-frame dynamics. Based on this finding, we propose DyMoS (Dynamic Motion Slider), a training-free and model-agnostic method that rebalances the attention pathway from generated frames to the reference frame during initial denoising steps. DyMoS leaves both the input image and model weights unchanged and introduces a single scalar parameter for continuous control over motion strength. Experiments across multiple state-of-the-art I2V backbones demonstrate that DyMoS consistently improves motion dynamics while maintaining visual quality and fidelity to the reference image.

图像生成视频运动增强注意力机制无训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。