首个面向流式音视频生成的综合评测基准,解决实时交互与长时稳定难题。
StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation

- 构建分进度与交互双赛道评估框架,覆盖指令遵循与状态保持。
- 13个主流模型在长时生成中均出现时间漂移,交互响应存在延迟瓶颈。
- 适用于研究流式音视频生成、多模态交互系统的研究者与开发者。
生成模型的最新进展正推动视频生成迈向无界流式音视频生成,以实现真实世界的实时交互。然而,现有评测基准主要评估完整序列,难以捕捉流式特性。为此,我们提出首个专为流式音视频生成设计的综合性评测基准——StreamAV-Bench。该基准建立统一评估框架,包含进度跟踪(用于指令遵循与长时稳定性)和交互跟踪(用于交互响应与状态保留复用)。基于32个细粒度维度的专家验证案例,我们对13个代表性系统进行了全面评估。分析表明,当前模型在进度生成中普遍存在时间漂移问题,在交互控制中存在响应瓶颈。基于全面的失败分析,我们分享了推动原生联合音视频流模型发展的关键洞察。
原文摘要 · Abstract (English)
Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the first comprehensive benchmark tailored for streaming audio-video generation. StreamAV-Bench establishes a unified evaluation framework, including the progressive track for instruction adherence and long-horizon stability, and the interactive track for interactive response and state retention and reuse. With expert-verified evaluation cases across 32 fine-grained dimensions, we conduct an extensive evaluation of 13 representative systems. Our analysis reveals that current models suffer from temporal drift in progressive generation and responsiveness bottlenecks during interactive control. Based on a comprehensive failure analysis, we share insights to advance the development of native joint audio-video streaming models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。