将情绪控制从视频生成中分离,实现更精准的情感表达。
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

- 拆分情绪氛围、语义线索和时间动态,分别独立控制
- 情绪对齐提升最高达37%,时序波动降低48%
- 支持多情感类别、跨模型通用,无需重训练
情绪影响观众对场景的解读,但现有视频生成模型将全局氛围、情绪相关语义线索和时间演变过程耦合在单一文本条件中。本文提出EmoWorld框架,通过冻结的流匹配视频扩散变压器(Video DiT)解耦这些因素。预处理阶段从保持几何结构的中性与情绪编辑全景图中提取层特定的情绪方向和可复用线索库。推理时,视觉氛围调控(VAS)将氛围方向注入隐藏状态,语义情绪调控(SAS)分离出可独立缩放的提示残差以表示语义线索,时间情绪调控(TAS)在去噪过程和视频时间上插值端点残差场。在Wan2.2数据集上,VAS使目标情绪对齐提升19%,时序波动指标降低48%;SAS使目标情绪对齐提升37%,检测到的情绪相关线索增加36%;TAS使过渡单调性提升15%。EmoWorld在27种情绪类别下评估,覆盖文本到视频与图像到视频任务,支持多种Video-DiT骨干网络,且在相机条件构图中无需更新生成器参数。
原文摘要 · Abstract (English)
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。