让人脸动画精准控制眼神细节,支持思考、犯困等非表情状态
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

- 用三个AI代理分步解析高层指令生成面部关键点
- 在HDTF和EMH数据集上眼区控制精度优于现有方法
- 适合需要精细眼神动作的影视动画与虚拟角色开发
人像动画虽已实现高质量视觉效果与口型同步,但在眼部区域的细粒度控制仍面临输入粒度与运动精度的权衡。现有方法依赖情绪标签或粗略文本提示,难以描述细微眼动;而基于动作单元(AU)或驱动视频的方法虽精度高,但输入负担重。这些限制尤其影响对非情绪状态(如思考、困倦)的表达。为此,我们提出CogPortrait,一种两阶段框架,从高层标签生成人像动画。第一阶段,三个链式思维多模态大模型(MLLM)代理通过时间事件规划、原型检索与合成、语义-生理约束强制,将高层标签转化为面部关键点;第二阶段,基于DiT的视频生成主干在关键点、参考肖像、音频及文本提示条件下生成动画,结合动态无分类器引导策略与眼区感知重加权,以及基于KTO的边界情况优化。我们还引入了EMH基准,涵盖多样情绪与非情绪类别,并采用两个AU级指标评估眼区与头部运动的细粒度控制能力。在HDTF与EMH基准上的大量实验表明,CogPortrait在保持优异视觉质量与身份一致性的同时,实现了更精确的眼部区域控制。
原文摘要 · Abstract (English)
Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input granularity and motion accuracy. Existing methods using emotion labels or coarse text prompts are insufficient for describing subtle ocular dynamics, whereas approaches based on Action Units or driving videos provide higher fidelity at the cost of a heavier input burden. These limitations are still restrictive for beyond-emotion states (e.g., thinking) and drowsiness. In light of the above, we propose CogPortrait, a two-stage framework that generates portrait animations from high-level labels. In the first stage, three chain-of-thought Multimodal Large Language Models (MLLMs) agents compile high-level labels into facial keypoints through temporal event planning, prototype retrieval, and composition from a real-behavior library, and semantic-physiological constraint enforcement. In the second stage, a DiT-based video generation backbone synthesizes the final animation conditioned on the keypoints, reference portrait, audio, and text prompt, enhanced by a dynamic classifier-free guidance strategy with eye-region-aware reweighting and KTO-based refinement for boundary cases. We further introduce the EMH benchmark covering diverse emotions and beyond-emotion categories with two AU-level metrics for evaluating fine-grained eye-region and head-motion control. Extensive experiments on HDTF and the EMH benchmark demonstrate that CogPortrait achieves more precise eye-region control than existing methods while maintaining supe- rior visual quality and identity consistency
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。