arXiv:2504.10044cs.CV2025-04被引 8

用真人反馈提升动漫视频生成质量,解决画面扭曲和不连贯问题。

Aligning Anime Video Generation with Human Feedback

  • 构建首个多维度动漫视频奖励数据集,含3万条人工标注样本。
  • 提出AnimeReward模型,通过视觉-语言模型分维度评估画风与连贯性。
  • 引入GAPO方法显式优化偏好差距,提升生成效果与训练效率。

动漫视频生成受限于动漫数据稀缺和异常运动模式,常出现动作扭曲与闪烁伪影,导致与人类偏好不符。现有奖励模型主要针对真实世界视频设计,难以捕捉动漫特有的视觉风格与一致性需求。本文提出一个利用人类反馈增强动漫视频生成的流水线:首先构建首个多维度动漫视频奖励数据集,包含3万条人工标注样本,涵盖对视觉外观与视觉一致性的偏好;在此基础上,开发AnimeReward模型,采用专用视觉-语言模型分别评估不同评价维度以指导偏好对齐;进一步提出间隙感知偏好优化(GAPO)方法,将偏好差距显式融入优化过程,提升对齐性能与效率。大量实验表明,AnimeReward优于现有模型,GAPO在定量指标与人工评估中均表现更优,验证了该流水线在提升动漫视频质量上的有效性。代码与数据集已公开于https://github.com/bilibili/Index-anisora。

原文摘要 · Abstract (English)

Anime video generation faces significant challenges due to the scarcity of anime data and unusual motion patterns, leading to issues such as motion distortion and flickering artifacts, which result in misalignment with human preferences. Existing reward models, designed primarily for real-world videos, fail to capture the unique appearance and consistency requirements of anime. In this work, we propose a pipeline to enhance anime video generation by leveraging human feedback for better alignment. Specifically, we construct the first multi-dimensional reward dataset for anime videos, comprising 30k human-annotated samples that incorporating human preferences for both visual appearance and visual consistency. Based on this, we develop AnimeReward, a powerful reward model that employs specialized vision-language models for different evaluation dimensions to guide preference alignment. Furthermore, we introduce Gap-Aware Preference Optimization (GAPO), a novel training method that explicitly incorporates preference gaps into the optimization process, enhancing alignment performance and efficiency. Extensive experiment results show that AnimeReward outperforms existing reward models, and the inclusion of GAPO leads to superior alignment in both quantitative benchmarks and human evaluations, demonstrating the effectiveness of our pipeline in enhancing anime video quality. Our code and dataset are publicly available at https://github.com/bilibili/Index-anisora.

动漫生成奖励模型人类反馈视频质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。