arXiv:2606.13304cs.CV2026-06

让说话人物视频更自然,同时精准对口型且表情生动。

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance

论文配图:ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance
图 1 · 摘自论文原文
  • 用多层级语音表征引导生成,兼顾发音细节与面部表情
  • 无需人工标注或奖励模型,通过无奖赏强化学习优化头部动作
  • 在口型同步和表达自然度上均超越现有方法

语音驱动的说话人物动画旨在生成逼真的肖像视频,使面部动作与语音自然同步。尽管视频生成技术进步显著,但实现精确的唇形同步与丰富的情感表达仍具挑战。现有方法常在准确性与表现力间权衡,导致动画或僵硬或不同步。我们提出ReFree-S2V,一种基于流匹配的语音到肖像动画框架,利用预训练视频生成模型,融合多层次语音表征(包含音素与语调信息)以实现细粒度发音控制与高层情感提示。该模型通过可学习的选择器将不同层次的语音信息注入Transformer模块,兼顾唇形同步与自然表情运动。为生成自然头部动作,引入新型无奖赏强化学习机制,在不依赖人工设计的同步指标或奖励模型的前提下,避免不合理的运动。大量实验表明,ReFree-S2V在定量口型同步准确率与定性人类评估中均达到领先水平。

原文摘要 · Abstract (English)

Speech-driven talking character animation seeks to generate life-like portrait videos that convey natural conversation behavior, aligning facial motion with spoken audio. Although recent advances in video generation have substantially improved realism in video-based animation, achieving both accurate lip articulation and expressive behavior remains challenging. Existing approaches typically trade off precise phoneme-to-lip synchronization against dynamic facial expressions and head motion, yielding animations that are either accurate yet rigid, or expressive but poorly synchronized. We address this challenge by proposing ReFree-S2V, a flow-matching speech-to-portrait animation framework that builds upon a pretrained video generation model to achieve fine-grained speech articulation and high-level expressive cues in speech-driven portrait animation. This model introduces a multi-level speech representation capturing phonetic and prosodic information at both local and global granularities. These representations are selectively injected into transformer blocks via learnable level selectors, enabling both accurate lip synchronization and natural expressive motion. To achieve natural head movements, we further introduce a novel reward-free reinforcement learning scheme into flow-matching training to discourage perceptually implausible motion without relying on handcrafted synchronization metrics or reward models, or the high cost of human preference annotation. Extensive experiments demonstrate that ReFree-S2V achieves state-of-the-art performance, significantly outperforming existing methods in both quantitative lip-sync accuracy and qualitative human evaluations of naturalness and expressivity.

视频生成语音驱动表情自然强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。