arXiv:2603.14794cs.CVcs.LG2026-03被引 1

构建70小时双人对话数据集,用于研究实时互动行为建模。

Face-to-Face: A Video Dataset for Multi-Person Interaction Modeling

  • 采用半自动流程提取双人对话的时序对齐视频与音频轨道。
  • 在跨人物视觉上下文中生成数字虚拟人,提升情感一致性和流畅性。
  • 适合研究对话互动、多主体行为建模的研究者使用。

现有音视频数据集多为孤立说话人短独白,难以刻画真实对话中的反应节奏。本文提出《Face-to-Face with Jimmy Fallon (F2F-JF)》,一个70小时、包含14,000个片段的双人脱口秀对话数据集,保留嘉宾发言与主持人回应间的时序依赖关系。通过多主体跟踪、语音分段与轻量人工验证的半自动流程,提取出精准裁剪的主客双方视频轨道及元数据,可直接用于下游建模。以多说话人扩散模型为基础,基于前一说话人视频和当前说话人音频生成主持人视频,实现小而稳定的表情-身份一致性(Emotion-FID)与视频质量(FVD)提升,同时保持唇动同步。数据集、处理流程与基线代码已公开,为研究双人序列互动行为提供完整解决方案。

原文摘要 · Abstract (English)

Modeling the reactive tempo of human conversation remains difficult because most audio-visual datasets portray isolated speakers delivering short monologues. We introduce \textbf{Face-to-Face with Jimmy Fallon (F2F-JF)}, a 70-hour, 14k-clip dataset of two-person talk-show exchanges that preserves the sequential dependency between a guest turn and the host's response. A semi-automatic pipeline combines multi-person tracking, speech diarization, and lightweight human verification to extract temporally aligned host/guest tracks with tight crops and metadata that are ready for downstream modeling. We showcase the dataset with a reactive, speech-driven digital avatar task in which the host video during $[t_1,t_2]$ is generated from their audio plus the guest's preceding video during $[t_0,t_1]$. Conditioning a MultiTalk-style diffusion model on this cross-person visual context yields small but consistent Emotion-FID and FVD gains while preserving lip-sync quality relative to an audio-only baseline. The dataset, preprocessing recipe, and baseline together provide an end-to-end blueprint for studying dyadic, sequential behavior, which we expand upon throughout the paper. Dataset and code are available at https://face2face2026.github.io.

对话建模视频生成多主体交互数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。