arXiv:2512.14938cs.CVcs.AI2025-12被引 1

开源6.3小时高分辨率语音驱动视频数据集,支持分钟级生成与低成本推理。

TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation

  • 基于2.3万段音视频同步素材构建,采用多阶段筛选与标注流程
  • 50亿参数模型实现分钟级生成,推理成本仅为140亿模型的十分之一
  • 支持零样本配音与智能剧本重写,适合视频生成与多模态研究者

我们提出TalkVerse,一个大规模、开放的单人语音驱动说话视频生成语料库,旨在实现方法间的公平、可复现对比。现有顶尖系统依赖封闭数据或计算密集型模型,而TalkVerse提供230万条高分辨率(720p/1080p)音视频同步片段,总计6.3千小时,源自超过6万小时视频,经场景切分检测、美学评估、严格音画同步校验及全面标注(含2D骨骼与结构化视觉/音频风格描述)处理。基于此,我们构建了基于Wan2.2-5B的可复现50亿参数DiT基线模型。通过使用高下采样比视频VAE与带运动帧上下文的滑动窗口机制,模型实现低漂移的分钟级生成,其唇形同步与视觉质量媲美140亿参数的Wan-S2V模型,但推理成本降低10倍。为增强长视频叙事能力,引入多模态大模型导演,根据音视频线索重写提示词。此外,模型支持通过受控潜在噪声注入实现零样本视频配音。项目已开源数据集、训练方案与50亿参数检查点,降低语音驱动人类视频生成研究门槛。

原文摘要 · Abstract (English)

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on closed data or compute-heavy models, TalkVerse offers 2.3 million high-resolution (720p/1080p) audio-video synchronized clips totaling 6.3k hours. These are curated from over 60k hours of video via a transparent pipeline that includes scene-cut detection, aesthetic assessment, strict audio-visual synchronization checks, and comprehensive annotations including 2D skeletons and structured visual/audio-style captions. Leveraging TalkVerse, we present a reproducible 5B DiT baseline built on Wan2.2-5B. By utilizing a video VAE with a high downsampling ratio and a sliding window mechanism with motion-frame context, our model achieves minute-long generation with low drift. It delivers comparable lip-sync and visual quality to the 14B Wan-S2V model but with 10$\times$ lower inference cost. To enhance storytelling in long videos, we integrate an MLLM director to rewrite prompts based on audio and visual cues. Furthermore, our model supports zero-shot video dubbing via controlled latent noise injection. We open-source the dataset, training recipes, and 5B checkpoints to lower barriers for research in audio-driven human video generation. Project Page: https://zhenzhiwang.github.io/talkverse/

视频生成语音驱动扩散模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。