arXiv:2607.11081cs.CVcs.AI2026-07中稿 · ECCV

通过注意力头控制扩散模型,实现精准动作迁移。

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

论文配图:Controlling Motion Transfer in Diffusion Transformers via Attention Heads
图 1 · 摘自论文原文
  • 按注意力头区分运动与结构特征,针对性提取动作信息。
  • 无需更新参数,仅靠特征重调即可实现高保真动作迁移。
  • 方法可解释性强,适合需要精细动作控制的视频生成场景。

扩散变压器(DiTs)在视频生成中取得了高质量、时间连贯的结果。然而,将它们扩展到动作迁移任务仍具挑战性,因为对DiTs内部运动与结构表征的理解有限。本文从注意力头层面分析视频DiTs,发现部分头专门处理运动,另一些头则专注空间结构。基于此,我们提出一种无需参数更新的头感知可控动作迁移框架:通过语义对应引导优化运动头输出,并选择性注入结构特征以保持画面完整性。该方法不仅实现了精准的动作迁移,还为可调控视频生成提供了可解释的基础。

原文摘要 · Abstract (English)

Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.

扩散模型动作迁移注意力机制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。