通过重注意力机制实现文本引导的视频编辑精准定位。
Re-Attentional Controllable Video Diffusion Editing
- 在去噪阶段重调文本与视频的交叉注意力,精准对齐物体位置。
- 显著减少物体错位、数量错误等控制问题,提升编辑一致性。
- 适合需要高精度视频修改的研究者与创意工作者。
文本引导的视频编辑因流程简化而受到关注,用户仅需修改对应视频的文本提示即可完成编辑。现有研究利用大规模文生图扩散模型实现文本引导视频编辑,取得显著效果。然而仍存在物体位置错乱、数量错误等可控性挑战。为此,本文提出训练-free的重注意力可控视频扩散编辑方法(ReAtCo)。针对空间位置对齐问题,提出重注意力扩散(RAD),在去噪阶段重构文本提示与目标视频间的交叉注意力响应,实现空间位置对齐与语义高保真。针对不变区域内容保持问题,提出不变区域引导联合采样(IRJS)策略,缓解每步去噪中不变区域的固有采样误差,约束生成内容与不变区域一致。实验表明,ReAtCo持续提升视频扩散编辑的可控性,性能优于现有方法。
原文摘要 · Abstract (English)
Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resulting in remarkable video editing capabilities. However, they may still suffer from some limitations such as mislocated objects, incorrect number of objects. Therefore, the controllability of video editing remains a formidable challenge. In this paper, we aim to challenge the above limitations by proposing a Re-Attentional Controllable Video Diffusion Editing (ReAtCo) method. Specially, to align the spatial placement of the target objects with the edited text prompt in a training-free manner, we propose a Re-Attentional Diffusion (RAD) to refocus the cross-attention activation responses between the edited text prompt and the target video during the denoising stage, resulting in a spatially location-aligned and semantically high-fidelity manipulated video. In particular, to faithfully preserve the invariant region content with less border artifacts, we propose an Invariant Region-guided Joint Sampling (IRJS) strategy to mitigate the intrinsic sampling errors w.r.t the invariant regions at each denoising timestep and constrain the generated content to be harmonized with the invariant region content. Experimental results verify that ReAtCo consistently improves the controllability of video diffusion editing and achieves superior video editing performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。