无需训练即可精准控制多人视频中的互动行为。
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

- 通过时空掩码和方向重加权模块,动态调节注意力分布。
- 在多个模型上显著提升交互真实度,接近真人视频水平。
- 适合影视生成、社交机器人等需要自然互动的场景。
视频生成技术已能从文本或图像生成逼真视频,但影视制作与社交机器人对多人互动视频的需求日益增长,如对话、手势和协同动作。现有模型缺乏对互动行为的显式控制,常出现动作执行者错误、互动顺序混乱、目标指向偏差等问题。为此,我们提出SocialDirector,一种无需训练的交互控制方法,通过调制交叉注意力图来增强生成模型。其包含两个模块:社会角色掩码(Social Actor Masking)利用时空掩码使每个人只关注自身文本描述,避免动作-角色错配与互动紊乱;方向重加权(Directional Reweighting)强化对方向词(如'左向'、'右向')的注意力,确保动作朝向正确目标。为评估交互质量,我们在现有数据集上标注互动描述,并构建基于开源视觉语言模型的自动化评估流程。实验表明,SocialDirector在不同视频生成模型上均显著提升交互保真度,接近真人视频上限。
原文摘要 · Abstract (English)
Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including conversations, gestures, and coordinated actions. However, existing models offer no explicit control over interactions, such as who performs which action, when it occurs, and toward whom it is directed. This often results in wrong person performing unintended actions (actor-action mismatch), disordered social dynamics, and wrong action targets. To address these challenges, we present SocialDirector, a training-free interaction controller that enhances the generation model by modulating cross-attention maps. SocialDirector contains two modules: Social Actor Masking and Directional Reweighting. Social Actor Masking constrains each person's visual tokens to attend only to their own textual descriptions via a spatiotemporal mask, avoiding actor-action mismatch and disordered social dynamics. Directional Reweighting amplifies attention to directional words (e.g., "leftward", "right"), leading each action towards its intended target. To evaluate generated social interactions, we annotate existing datasets with interaction descriptions and build a fully automated evaluation pipeline powered by open-source VLMs. Experiments on different video generation models show that SocialDirector significantly improves interaction fidelity and approaches the upper bound set by real videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。