通过分阶段优化提升语音驱动人像动画的自然度与清晰度
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
- 按去噪时间步划分阶段,分别优化动作与视觉质量
- 推理时仅需30次去噪步骤,速度提升3.3倍且画质损失小
- 适合需要高效生成高质量人像动画的研究者和开发者
基于扩散模型的人体视频生成与动画技术虽取得显著进展,但动作自然性与视觉保真度之间仍存在权衡。为此,我们提出AlignHuman框架,结合偏好优化与分而治之的训练策略,联合优化这两项竞争目标。分析去噪过程发现:早期时间步主要决定动作动态,后期时间步可有效控制保真度与人体结构,即使跳过早期步骤亦可。基于此,我们提出时间步段偏好优化(TPO),引入两个专用LoRA模块作为专家对齐单元,分别在对应时间区间内激活,以增强动作自然性与视觉保真度。大量实验表明,AlignHuman在不显著降低生成质量的前提下,将推理所需的去噪次数(NFE)从100降至30,实现3.3倍加速。
原文摘要 · Abstract (English)
Recent advancements in human video generation and animation tasks, driven by diffusion models, have achieved significant progress. However, expressive and realistic human animation remains challenging due to the trade-off between motion naturalness and visual fidelity. To address this, we propose \textbf{AlignHuman}, a framework that combines Preference Optimization as a post-training technique with a divide-and-conquer training strategy to jointly optimize these competing objectives. Our key insight stems from an analysis of the denoising process across timesteps: (1) early denoising timesteps primarily control motion dynamics, while (2) fidelity and human structure can be effectively managed by later timesteps, even if early steps are skipped. Building on this observation, we propose timestep-segment preference optimization (TPO) and introduce two specialized LoRAs as expert alignment modules, each targeting a specific dimension in its corresponding timestep interval. The LoRAs are trained using their respective preference data and activated in the corresponding intervals during inference to enhance motion naturalness and fidelity. Extensive experiments demonstrate that AlignHuman improves strong baselines and reduces NFEs during inference, achieving a 3.3$\times$ speedup (from 100 NFEs to 30 NFEs) with minimal impact on generation quality. Homepage: \href{https://alignhuman.github.io/}{https://alignhuman.github.io/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。