arXiv:2603.00159cs.CVcs.AI2026-03被引 2

用强化学习提升语音驱动人脸视频的唇同步与自然度。

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

论文配图:FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
图 1 · 摘自论文原文
  • 基于多模态自回归模型,结合强化学习优化生成过程。
  • 通过人类感知对齐评估系统提升唇同步与表情自然度。
  • 适合关注人脸动画质量与生成真实感的研究者。

由于唇同步不准确、动作不自然以及评价指标与人类感知相关性差,生成逼真说话人脸视频仍具挑战。我们提出 FlowPortrait,一种基于多模态骨干网络的强化学习框架,用于音频驱动的人脸视频生成。FlowPortrait 引入基于多模态大语言模型(MLLMs)的人类对齐评估系统,量化唇同步精度、表现力与运动质量。这些信号与感知一致性和时序一致性正则项结合,形成稳定复合奖励,通过分组相对策略优化(GRPO)对生成器进行后训练。大量实验包括自动评估与人工偏好研究均表明,FlowPortrait 持续生成更高质量的说话人脸视频,验证了强化学习在人脸动画中的有效性。

原文摘要 · Abstract (English)

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a reinforcement-learning framework for audio-driven portrait animation built on a multimodal backbone for autoregressive audio-to-video generation. FlowPortrait introduces a human-aligned evaluation system based on Multimodal Large Language Models (MLLMs) to assess lip-sync accuracy, expressiveness, and motion quality. These signals are combined with perceptual and temporal consistency regularizers to form a stable composite reward, which is used to post-train the generator via Group Relative Policy Optimization (GRPO). Extensive experiments, including both automatic evaluations and human preference studies, demonstrate that FlowPortrait consistently produces higher-quality talking-head videos, highlighting the effectiveness of reinforcement learning for portrait animation.

人脸生成强化学习语音驱动视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。