arXiv:2510.07249cs.CV2025-10NeurIPS被引 4

构建164万段多镜头演讲视频数据集,支持可控长视频生成

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

  • 构建多镜头、多视角的高质视频数据集,支持复杂运镜与动作建模
  • 基于164万段视频训练,显著提升生成视频的镜头连贯性与视觉吸引力
  • 适合做多镜头视频生成、跨模态建模的研究者使用

本文提出TalkCuts,一个大规模多镜头人类演讲视频数据集,旨在推动多镜头演讲视频生成研究。与以往聚焦单镜头、固定视角的数据集不同,TalkCuts包含164万段视频,总时长超过500小时,涵盖特写、半身和全身等多种镜头视角。数据集包含详细的文本描述、2D关键点及3D SMPL-X运动标注,覆盖超过10,000个身份,支持多模态学习与评估。作为首个展示该数据集价值的尝试,我们提出了Orator——一个基于大语言模型的多模态生成框架,其中语言模型充当多维度导演,协调相机切换、手势动作与语音变化。该架构通过集成多模态生成模块,实现长时序视频的连贯合成。在姿态引导与音频驱动设置下的大量实验表明,使用TalkCuts训练可显著提升生成视频的电影级连贯性与视觉表现力。我们相信TalkCuts为可控多镜头演讲视频生成及更广泛的多模态学习奠定了坚实基础。

原文摘要 · Abstract (English)

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality human speech videos with diverse camera shots, including close-up, half-body, and full-body views. The dataset includes detailed textual descriptions, 2D keypoints and 3D SMPL-X motion annotations, covering over 10k identities, enabling multimodal learning and evaluation. As a first attempt to showcase the value of the dataset, we present Orator, an LLM-guided multi-modal generation framework as a simple baseline, where the language model functions as a multi-faceted director, orchestrating detailed specifications for camera transitions, speaker gesticulations, and vocal modulation. This architecture enables the synthesis of coherent long-form videos through our integrated multi-modal video generation module. Extensive experiments in both pose-guided and audio-driven settings show that training on TalkCuts significantly enhances the cinematographic coherence and visual appeal of generated multi-shot speech videos. We believe TalkCuts provides a strong foundation for future work in controllable, multi-shot speech video generation and broader multimodal learning.

视频生成多镜头多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。