arXiv:2605.07545cs.CVcs.AI2026-05中稿 · ICML

无需配对数据,用自生成样本优化手部动作生成质量

Implicit Preference Alignment for Human Image Animation

论文配图:Implicit Preference Alignment for Human Image Animation
图 1 · 摘自论文原文
  • 通过自生成高质量样本最大化概率实现隐式偏好对齐
  • 手部感知局部优化使手部动作更自然,显著提升生成质量
  • 免去繁琐配对标注,降低数据构建成本,适合实用化部署

人体图像动画虽有显著进展,但手部动作因自由度高、运动复杂,仍难以生成高保真结果。传统基于人类反馈的强化学习需严格配对偏好数据,但动态手部区域帧间不一致导致标注成本过高。本文提出隐式偏好对齐(IPA),一种数据高效的后训练框架,无需配对偏好数据。其理论基础为隐式奖励最大化,通过最大化自生成高质量样本的概率并惩罚偏离预训练先验的偏差来实现对齐。进一步引入手部感知局部优化机制,显式引导对齐过程聚焦手部区域。实验表明,该方法有效提升手部生成质量,同时大幅降低偏好数据构建门槛。代码已开源。

原文摘要 · Abstract (English)

Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their high degrees of freedom and motion complexity. While reinforcement learning from human feedback, particularly direct preference optimization, offers a potential solution, it necessitates the construction of strict preference pairs. However, curating such pairs for dynamic hand regions is prohibitively expensive and often impractical due to frame-wise inconsistencies. In this paper, we propose Implicit Preference Alignment (IPA), a data-efficient post-training framework that eliminates the need for paired preference data. Theoretically grounded in implicit reward maximization, IPA aligns the model by maximizing the likelihood of self-generated high-quality samples while penalizing deviations from the pretrained prior. Furthermore, we introduce a Hand-Aware Local Optimization mechanism to explicitly steer the alignment process toward hand regions. Experiments demonstrate that our method achieves effective preference optimization to enhance hand generation quality, while significantly lowering the barrier for constructing preference data. Codes are released at https://github.com/mdswyz/IPA

图像动画手部生成偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。