让视频人物边说边动,对话自然同步
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- 用跨模态映射实现音视频双向交互生成
- 在四方面评测中均优于现有方法
- 适合做虚拟助手、数字人对话系统
本文旨在从文本和参考图像中联合生成互动视频与对话语音。为构建类人对话系统,现有研究分别探索了说话头生成和对话语音生成,但通常孤立进行,忽略了人类对话中紧密耦合的音视频交互特性。本文提出统一框架TAVID,可同步生成互动人脸与对话语音。通过运动映射器和说话者映射器,实现音频与视觉模态间互补信息的双向交换。我们在四个维度上评估系统:说话脸真实度、听头响应性、双人对话流畅性与语音质量。大量实验表明,该方法在所有方面均表现优异。
原文摘要 · Abstract (English)
The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or listening head generation as well as conversational speech generation. However, these works are typically studied in isolation, overlooking the multimodal nature of human conversation, which involves tightly coupled audio-visual interactions. In this paper, we introduce TAVID, a unified framework that generates both interactive faces and conversational speech in a synchronized manner. TAVID integrates face and speech generation pipelines through two cross-modal mappers (i.e., a motion mapper and a speaker mapper), which enable bidirectional exchange of complementary information between the audio and visual modalities. We evaluate our system across four dimensions: talking face realism, listening head responsiveness, dyadic interaction fluency, and speech quality. Extensive experiments demonstrate the effectiveness of our approach across all these aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。