arXiv:2603.16711cs.CV2026-03

无需训练即可精准控制图像生成视频中的物体运动

Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search

  • 基于首尾帧运动先验,实现物体重定位而不需标注轨迹
  • 新方法在无真实轨迹情况下仍保持高运动保真度
  • 适合需要快速编辑且无标注数据的视频生成场景

我们提出Search2Motion,一种无需训练的图像到视频生成中对象级运动编辑框架。与以往需轨迹、边界框、掩码或运动场的方法不同,Search2Motion采用目标帧控制,利用首尾帧运动先验实现物体重定位,并保持场景稳定,无需微调。通过语义引导的对象插入和鲁棒背景修复实现可靠的目标帧构建。我们进一步发现早期步自注意力图可预测物体与相机动态,提供可解释用户反馈,并由此提出轻量级搜索策略ACE-Seed(早期种子注意力一致性),提升运动保真度,无需前瞻采样或外部评估器。针对现有基准混淆物体与相机运动的问题,我们引入S2M-DAVIS和S2M-OMB用于固定摄像头下物体独立评估,以及FLF2V-obj指标分离物体伪影,无需真实轨迹。Search2Motion在FLF2V-obj和VBench上持续优于基线。

原文摘要 · Abstract (English)

We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks, or motion fields, Search2Motion adopts target-frame-based control, leveraging first-last-frame motion priors to realize object relocation while preserving scene stability without fine-tuning. Reliable target-frame construction is achieved through semantic-guided object insertion and robust background inpainting. We further show that early-step self-attention maps predict object and camera dynamics, offering interpretable user feedback and motivating ACE-Seed (Attention Consensus for Early-step Seed selection), a lightweight search strategy that improves motion fidelity without look-ahead sampling or external evaluators. Noting that existing benchmarks conflate object and camera motion, we introduce S2M-DAVIS and S2M-OMB for stable-camera, object-only evaluation, alongside FLF2V-obj metrics that isolate object artifacts without requiring ground-truth trajectories. Search2Motion consistently outperforms baselines on FLF2V-obj and VBench.

视频生成运动控制无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。