arXiv:2504.08366cs.GRcs.CV2025-04SIGGRAPH被引 11

用两张单视角图生成流畅4D动态,无需预设物体类型和运动形式。

In-2-4D: Inbetweening from Two Single-View Images to 4D Generation

  • 分层识别关键帧,结合高斯点云与变形场生成局部3D动态片段。
  • 通过自注意力扩展与刚性约束,提升时序一致性与3D运动精度。
  • 适合需要精确控制动作过渡的动画生成、数字人建模等场景。

我们提出新任务In-2-4D,即从两幅单视角图像生成4D(3D+运动)中间过程。不同于仅基于文本或单图的视频/4D生成,本方法通过插值实现更精准的动作控制。给定表示物体起始与结束状态的两幅单目RGB图像,目标是生成并重建无类别、运动类型、长度与复杂度假设下的4D运动。为处理任意多样化的运动,采用基础视频插值模型预测运动。针对大帧间运动带来的歧义,采用分层策略识别视觉接近输入状态但运动显著的关键帧,再生成平滑片段。对每个片段,利用高斯溅射(3DGS)构建关键帧的3D表示,时间帧引导其通过形变场转换为动态3DGS。为增强时序一致性并优化3D运动,扩展多视角扩散模型的自注意力至时间维度,并引入刚性变换正则化。最后通过插值边界形变场合并独立生成的3D运动段,并优化对齐引导视频,确保过渡平滑无闪烁。大量定性定量实验及用户研究验证了方法有效性与设计合理性。

原文摘要 · Abstract (English)

We pose a new problem, In-2-4D, for generative 4D (i.e., 3D + motion) inbetweening to interpolate two single-view images. In contrast to video/4D generation from only text or a single image, our interpolative task can leverage more precise motion control to better constrain the generation. Given two monocular RGB images representing the start and end states of an object in motion, our goal is to generate and reconstruct the motion in 4D, without making assumptions on the object category, motion type, length, or complexity. To handle such arbitrary and diverse motions, we utilize a foundational video interpolation model for motion prediction. However, large frame-to-frame motion gaps can lead to ambiguous interpretations. To this end, we employ a hierarchical approach to identify keyframes that are visually close to the input states while exhibiting significant motions, then generate smooth fragments between them. For each fragment, we construct a 3D representation of the keyframe using Gaussian Splatting (3DGS). The temporal frames within the fragment guide the motion, enabling their transformation into dynamic 3DGS through a deformation field. To improve temporal consistency and refine the 3D motion, we expand the self-attention of multi-view diffusion across timesteps and apply rigid transformation regularization. Finally, we merge the independently generated 3D motion segments by interpolating boundary deformation fields and optimizing them to align with the guiding video, ensuring smooth and flicker-free transitions. Through extensive qualitative and quantitive experiments as well as a user study, we demonstrate the effectiveness of our method and design choices.

4D生成图像插值高斯溅射动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。