arXiv:2603.23132cs.CV2026-03

让两人对话视频更自然,通过动作引导实现精准反应控制

InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance

  • 用参考视频提取通用动作先验,实现身份无关的视频重演
  • 通过多模态大模型解析语音意图,精确控制反应时机与合理性
  • 针对极端头部姿态优化口型同步,适合需要精细交互的生成任务

尽管语音到视频合成取得进展,现有方法在捕捉双人互动中的跨个体依赖关系及对反应行为的细粒度控制方面仍存在困难。为此,我们提出InterDyad框架,通过查询结构化运动引导实现自然交互动态合成。首先设计了互异性注入器(Interactivity Injector),基于参考视频提取的身份无关动作先验实现视频重演;在此基础上,引入基于元查询的模态对齐机制,弥合对话音频与动作先验之间的差距。利用多模态大语言模型(MLLM),框架可从音频中提炼语言意图,以决定反应的精确时间点与恰当性。为提升极端头部姿态下的唇音同步质量,提出角色感知的双人高斯引导(RoDG),增强口型同步与空间一致性。最后,构建专用评估套件,包含新设计的度量指标以量化双人互动效果。大量实验表明,InterDyad显著优于现有最优方法,在生成自然且情境相关的双人互动视频方面表现卓越。

原文摘要 · Abstract (English)

Despite progress in speech-to-video synthesis, existing methods often struggle to capture cross-individual dependencies and provide fine-grained control over reactive behaviors in dyadic settings. To address these challenges, we propose InterDyad, a framework that enables naturalistic interactive dynamics synthesis via querying structural motion guidance. Specifically, we first design an Interactivity Injector that achieves video reenactment based on identity-agnostic motion priors extracted from reference videos. Building upon this, we introduce a MetaQuery-based modality alignment mechanism to bridge the gap between conversational audio and these motion priors. By leveraging a Multimodal Large Language Model (MLLM), our framework is able to distill linguistic intent from audio to dictate the precise timing and appropriateness of reactions. To further improve lip-sync quality under extreme head poses, we propose Role-aware Dyadic Gaussian Guidance (RoDG) for enhanced lip-synchronization and spatial consistency. Finally, we introduce a dedicated evaluation suite with novelly designed metrics to quantify dyadic interaction. Comprehensive experiments demonstrate that InterDyad significantly outperforms state-of-the-art methods in producing natural and contextually grounded two-person interactions. Please refer to our project page for demo videos: https://interdyad.github.io/.

语音生成视频双人交互动作控制唇同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。