arXiv:2511.14848cs.CV2025-11SIGGRAPH被引 3

无需绑定骨骼,跨类别物体也能精准传递3D动作。

Gaussian See, Gaussian Do: Semantic 3D Motion Transfer from Multiview Video

  • 用多视角视频反推动作嵌入,驱动目标形状生成动态画面。
  • 在真实视频上实现高保真动作迁移,结构一致性优于现有方法。
  • 适合做3D角色动画、虚拟试衣等需要灵活动作迁移的场景。

我们提出Gaussian See, Gaussian Do,一种从多视角视频进行语义3D动作迁移的新方法。该方法实现无绑定骨骼、跨类别物体间的语义对齐动作转移。基于隐式动作迁移技术,通过条件反演从源视频中提取动作嵌入,应用于静态目标形状的渲染帧,并利用生成的视频监督动态3D高斯点云重建。提出基于锚点的视图感知动作嵌入机制,确保跨视角一致性并加速收敛;设计鲁棒的4D重建流程,整合噪声监督视频。建立首个语义3D动作迁移基准,相比适配基线展现出更优的动作保真度与结构一致性。代码与数据见 https://gsgd-motiontransfer.github.io/

原文摘要 · Abstract (English)

We present Gaussian See, Gaussian Do, a novel approach for semantic 3D motion transfer from multiview video. Our method enables rig-free, cross-category motion transfer between objects with semantically meaningful correspondence. Building on implicit motion transfer techniques, we extract motion embeddings from source videos via condition inversion, apply them to rendered frames of static target shapes, and use the resulting videos to supervise dynamic 3D Gaussian Splatting reconstruction. Our approach introduces an anchor-based view-aware motion embedding mechanism, ensuring cross-view consistency and accelerating convergence, along with a robust 4D reconstruction pipeline that consolidates noisy supervision videos. We establish the first benchmark for semantic 3D motion transfer and demonstrate superior motion fidelity and structural consistency compared to adapted baselines. Code and data for this paper available at https://gsgd-motiontransfer.github.io/

3D动作迁移高斯溅射多视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。