arXiv:2509.06723cs.CV2025-09AAAI被引 6

零样本视频生成新方法,让动作指令更真实自然。

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

  • 引入3D感知运动投影,纠正视角偏差。
  • 测试时训练动态优化局部模型,保持生成质量。
  • 适合需要精准动作控制的视觉创作人群。

轨迹引导图像到视频生成旨在根据用户指定的动作指令合成视频。现有方法通常依赖计算成本高昂的微调和稀缺标注数据集。尽管一些零样本方法尝试在潜在空间中进行轨迹控制,但可能因忽略3D透视而产生不真实运动,并导致操纵后的潜在表示与网络噪声预测不一致。为此,我们提出Zo3T,一种新型零样本测试时训练框架,具备三项核心创新:首先,引入3D感知运动投影,通过推断场景深度获取视角正确的仿射变换以作用于目标区域;其次,提出轨迹引导测试时LoRA机制,动态注入并优化临时LoRA适配器至去噪网络,伴随潜在状态变化,由区域特征一致性损失驱动协同适应,有效施加运动约束的同时允许预训练模型局部调整内部表示,从而保证生成保真度与流形一致性;最后,设计引导场校正模块,通过一步前瞻策略优化条件引导场,修正去噪演化路径,确保高效生成向目标轨迹收敛。Zo3T显著提升了轨迹控制下的3D真实感与运动准确性,在性能上超越现有训练依赖及零样本方法。

原文摘要 · Abstract (English)

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the latent space, they may yield unrealistic motion by neglecting 3D perspective and creating a misalignment between the manipulated latents and the network's noise predictions. To address these challenges, we introduce Zo3T, a novel zero-shot test-time-training framework for trajectory-guided generation with three core innovations: First, we incorporate a 3D-Aware Kinematic Projection, leveraging inferring scene depth to derive perspective-correct affine transformations for target regions. Second, we introduce Trajectory-Guided Test-Time LoRA, a mechanism that dynamically injects and optimizes ephemeral LoRA adapters into the denoising network alongside the latent state. Driven by a regional feature consistency loss, this co-adaptation effectively enforces motion constraints while allowing the pre-trained model to locally adapt its internal representations to the manipulated latent, thereby ensuring generative fidelity and on-manifold adherence. Finally, we develop Guidance Field Rectification, which refines the denoising evolutionary path by optimizing the conditional guidance field through a one-step lookahead strategy, ensuring efficient generative progression towards the target trajectory. Zo3T significantly enhances 3D realism and motion accuracy in trajectory-controlled I2V generation, demonstrating superior performance over existing training-based and zero-shot approaches.

视频生成零样本3D感知测试时训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。