用姿态引导生成更清晰、连贯的手语视频,解决细节模糊和帧间跳变问题。
Pose-Guided Fine-Grained Sign Language Video Generation
- 通过光流变形与姿态融合,分步实现粗粒度运动与细粒度生成
- 在多个基准测试中优于现有方法,细节更清晰、时间一致性更强
- 适合手语生成、无障碍传播及影视动画领域研究者参考
手语视频是推广和学习手语的重要媒介。然而,现有大多数人体图像合成方法生成的手语图像存在细节失真、模糊或结构错误的问题,且视频帧之间缺乏时间一致性,常出现闪烁或前后帧细节突变等异常。为此,我们提出一种新的姿态引导运动模型(PGMM),用于生成精细且运动连贯的手语视频。首先,设计粗粒度运动模块(CMM),通过光流变形完成特征形变,实现粗结构运动迁移而不改变外观;其次,提出姿态融合模块(PFM),引导RGB与姿态特征的模态融合,完成细粒度生成;最后,设计时间一致性差异(TCD)新指标,通过比较重建视频帧与目标视频前后帧的差异,定量评估视频的时间一致性。大量定性和定量实验表明,本方法在多数基准测试中优于当前最先进方法,显著提升细节表现与时间一致性。
原文摘要 · Abstract (English)
Sign language videos are an important medium for spreading and learning sign language. However, most existing human image synthesis methods produce sign language images with details that are distorted, blurred, or structurally incorrect. They also produce sign language video frames with poor temporal consistency, with anomalies such as flickering and abrupt detail changes between the previous and next frames. To address these limitations, we propose a novel Pose-Guided Motion Model (PGMM) for generating fine-grained and motion-consistent sign language videos. Firstly, we propose a new Coarse Motion Module (CMM), which completes the deformation of features by optical flow warping, thus transfering the motion of coarse-grained structures without changing the appearance; Secondly, we propose a new Pose Fusion Module (PFM), which guides the modal fusion of RGB and pose features, thus completing the fine-grained generation. Finally, we design a new metric, Temporal Consistency Difference (TCD) to quantitatively assess the degree of temporal consistency of a video by comparing the difference between the frames of the reconstructed video and the previous and next frames of the target video. Extensive qualitative and quantitative experiments show that our method outperforms state-of-the-art methods in most benchmark tests, with visible improvements in details and temporal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。