用手机自拍两张照+动作数据,就能生成逼真人像视频。
Generating Fit Check Videos with a Handheld Camera
- 输入前后两张自拍和手机运动数据,合成连贯人像视频。
- 生成视频含一致光照与阴影,支持新场景渲染。
- 创新扩散模型设计,提升细节真实感,适合普通用户。
自拍全身视频流行,但多数需固定摄像头、精心构图和反复练习。本文提出一种更便捷方案:仅需手持手机拍摄两幅镜中正背照及动作轨迹(IMU),即可合成你完成目标动作的逼真视频。支持将人物渲染至新场景,保持光照与阴影一致性。我们提出基于视频扩散的新模型,采用无参数帧生成策略和多参考注意力机制,有效融合前后视角外观信息;并引入图像微调策略,显著提升画面清晰度,改善阴影与反射效果,增强人-景融合的真实感。
原文摘要 · Abstract (English)
Self-captured full-body videos are popular, but most deployments require mounted cameras, carefully-framed shots, and repeated practice. We propose a more convenient solution that enables full-body video capture using handheld mobile devices. Our approach takes as input two static photos (front and back) of you in a mirror, along with an IMU motion reference that you perform while holding your mobile phone, and synthesizes a realistic video of you performing a similar target motion. We enable rendering into a new scene, with consistent illumination and shadows. We propose a novel video diffusion-based model to achieve this. Specifically, we propose a parameter-free frame generation strategy and a multi-reference attention mechanism to effectively integrate appearance information from both the front and back selfies into the video diffusion model. Further, we introduce an image-based fine-tuning strategy to enhance frame sharpness and improve shadows and reflections generation for more realistic human-scene composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。