不依赖音频,精准对齐口型与语音,提升字词级和音素级同步精度。
Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning
- 结合局部上下文感知特征提取与多任务学习,增强对细微口型变化的敏感度。
- 在LRS2数据集上,词级准确率提升6%,音素级提升27%。
- 适合自动字幕生成、短视频平台内容标注等应用场景。
本文提出一种新型视觉强制对齐(VFA)方法,无需依赖音频信号即可精准同步语音与对应口型动作。该方法融合局部上下文感知特征提取器,并通过多任务学习优化全局与局部上下文特征,提升对微小唇部运动的敏感性,实现更精确的词级与音素级对齐。引入改进的维特比算法进行后处理,显著减少错位。实验结果表明,在LRS2数据集上,该方法相比现有方法在词级准确率上提升6%,音素级提升27%,为电视剧自动字幕生成或TikTok、YouTube Shorts等用户生成内容平台的智能标注提供了新可能。
原文摘要 · Abstract (English)
This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multi-task learning to refine both global and local context features, enhancing sensitivity to subtle lip movements for precise word-level and phoneme-level alignment. Incorporating the improved Viterbi algorithm for post-processing, our method significantly reduces misalignments. Experimental results show our approach outperforms existing methods, achieving a 6% accuracy improvement at the word-level and 27% improvement at the phoneme-level in LRS2 dataset. These improvements offer new potential for applications in automatically subtitling TV shows or user-generated content platforms like TikTok and YouTube Shorts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。