让动画视频生成摆脱物理真实,用艺术规则驱动创作。
AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics

- 用生产知识体系提取动画风格、运镜等可控变量,实现精准导演指令注入。
- 三阶段训练:重定义正确性、打破物理先验、区分艺术表达与生成崩溃。
- 专业动画师评测中四项领先,尤其在理解提示词和艺术化运动上提升显著。
视频生成模型通常以物理真实性为先验,但动漫刻意违背物理规律:如拖影、撞击帧、角色变形等,且存在数千种共存的艺术惯例,无法归纳出统一的“动漫物理法则”。现有模型因此要么弱化艺术特征,要么在风格差异下崩溃。我们提出AniMatrix,通过双通道条件机制,聚焦艺术而非物理正确性。首先,生产知识系统将动漫结构化为可控制的生产变量(风格、运动、镜头、特效),AniCaption从图像中推断这些变量作为导演指令;可训练标签编码器保持字段-值结构,冻结的T5处理自由文本;双路径注入(交叉注意力精细控制,AdaLN调制全局强化)确保类别指令不被文本稀释。其次,采用风格-运动-形变课程学习,逐步从近似物理运动过渡到完全动漫表现力。第三,使用领域特定奖励模型进行形变感知偏好优化,区分有意的艺术表达与病理式崩溃。在由专业动画师评估的动漫专属人类评测中,AniMatrix在五项生产维度中四项排名第一,相比Seedance-Pro 1.0在提示理解(+0.70,+22.4%)和艺术化运动(+0.55,+16.9%)上提升最大。配套资源即将公开,支持可复现与后续研究。
原文摘要 · Abstract (English)
Video generation models internalize physical realism as their prior. Anime deliberately violates physics: smears, impact frames, chibi shifts; and its thousands of coexisting artistic conventions yield no single "physics of anime" a model can absorb. Physics-biased models therefore flatten the artistry that defines the medium or collapse under its stylistic variance. We present AniMatrix, a video generation model that targets artistic rather than physical correctness through a dual-channel conditioning mechanism and a three-step transition: redefine correctness, override the physics prior, and distinguish art from failure. First, a Production Knowledge System encodes anime as a structured taxonomy of controllable production variables (Style, Motion, Camera, VFX), and AniCaption infers these variables from pixels as directorial directives. A trainable tag encoder preserves the field-value structure of this taxonomy while a frozen T5 encoder handles free-form narrative; dual-path injection (cross-attention for fine-grained control, AdaLN modulation for global enforcement) ensures categorical directives are never diluted by open-ended text. Second, a style-motion-deformation curriculum transitions the model from near-physical motion to full anime expressiveness. Third, deformation-aware preference optimization with a domain-specific reward model separates intentional artistry from pathological collapse. On an anime-specific human evaluation with five production dimensions scored by professional animators, AniMatrix ranks first on four of five, with the largest gains over Seedance-Pro 1.0 on Prompt Understanding (+0.70, +22.4 percent) and Artistic Motion (+0.55, +16.9 percent). We are preparing accompanying resources for public release to support reproducibility and follow-up research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。