用时间轴精准控制人脸动作生成,实现毫秒级动作对齐。
Exploring Timeline Control for Facial Motion Generation
- 通过时间轴多轨标注,实现帧级精细动作时序控制。
- 基于扩散模型生成动作自然且严格对齐输入时间轴。
- 支持文本转时间轴,适合动画、虚拟人等高精度场景。
本文提出一种新的面部动作生成控制信号——时间轴控制。相比音频和文本信号,时间轴能提供更精细的控制,例如精确生成特定面部动作及其时间点。用户可指定多轨时间轴,按时间间隔排列面部动作,实现每个动作的精准时序控制。为建模时间轴控制能力,我们首先在自然面部动作序列中以帧级粒度标注动作时间区间,采用基于Toeplitz逆协方差聚类的方法降低人工标注成本。基于标注数据,提出一种基于扩散模型的生成方法,能够生成自然且与输入时间轴高度对齐的面部动作。该方法还支持通过ChatGPT将文本转化为时间轴,实现文本引导生成。实验表明,本方法能以较高准确率标注动作区间,并生成与时间轴精确对齐的自然面部动作。
原文摘要 · Abstract (English)
This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。