构建首个大规模交互式音视频响应数据集,让AI能真实回应用户互动。
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

- 从直播弹幕视频中提取真实交互片段,构建上下文-刺激-响应三元组。
- 涵盖超45万条交互数据,覆盖对话、物体、操作等五类场景。
- 适合研究多模态交互生成的开发者与研究员使用。
大语言模型使文本成为人机交互的默认媒介,但仅靠文本无法满足多模态助手、虚拟形象和具身智能体所需的完整响应。尽管近期音视频生成模型可合成高质量同步内容,但现有训练数据多为描述性标注,而非由外部用户互动触发的真实响应。我们提出「InteracVid」,首个开源的大规模交互式音视频响应数据集,每个样本均包含前序音视频上下文、外部刺激及随之产生的真实互动响应。设计了一种元数据感知的流水线,从长达数千小时、噪声密集的直播流中提取交互片段,共获得超过45.4万条上下文-查询-响应三元组,来自5.9万余段直播视频,覆盖以对话为中心、物体为中心、流程性、具身化和屏幕交互等五类场景。十名评估者的人工研究证实,提取的互动具有因果性、自然性和时间完整性,无论真实或重构的查询均成立。在100个真实直播弹幕查询的保留基准测试中,基于InteracVid微调显著提升了交互规划与音视频生成效果,独立人工评估复现了系统排名及自动评测的结论。结果凸显结构化交互数据对交互式多模态生成的关键作用。
原文摘要 · Abstract (English)
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。