用一句话文本生成真实感对话视频,解决真人对话数据难收集问题。
Conversational Human Audio-visual Talking Dialogue Generation

- 输入文本即可生成语音与表情同步的双人对话视频
- 在客观和主观评测中均优于现有方法,生成5万条高质量数据
- 适合数字人、虚拟助手开发,降低数据采集成本
大规模双人交互式音视频对话(DIAD)数据集是开发类人互动虚拟代理和数字人的基础资源,但其采集耗时、昂贵且存在伦理风险。为此,我们提出CHAT框架,仅需单个文本提示即可生成多样、成对且互响应的语音-面部对话片段。CHAT融合大语言模型与说话人脸模型,并引入交互式音频与面部行为优化模块,实现内容多样、身份多样的对偶对话视频生成。实验表明,CHAT在客观与主观评价上均优于现有同类方法。此外,我们构建的合成数据集CHAT-AVD-50k可作为下游交互头部生成的有效预训练数据,在REACT 2024上持续提升PerFRDiff与ReactDiff指标。CHAT为替代高成本、高伦理风险的真实数据采集提供了可扩展方案。
原文摘要 · Abstract (English)
Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. However, collecting such data is time-consuming, expensive, and ethically sensitive. To address this, we propose CHAT, a new dyadic interactive audio-visual dialogue generation (DIADG) framework that generates diverse, paired, and mutually responsive speech-face dialogue clips from a single textual prompt. CHAT unifies large language models and talking face models with interactive audio and facial behaviour refinement modules, enabling the generation of aligned dyadic dialogue clips with diverse contents and facial identities. Experiments show that CHAT outperforms existing related methods designed for similar tasks under both objective and subjective evaluations. Moreover, our synthesised CHAT-AVD-50k dataset serves as effective pre-training data for downstream interactive head generation, consistently improving PerFRDiff and ReactDiff on REACT 2024. CHAT offers a scalable alternative to the costly and ethically sensitive collection of real dyadic interaction data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。