arXiv:2605.00873cs.MMcs.AI2026-05

首个评估文本生成视频在不合理场景下表现的可靠基准

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

论文配图:BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios
图 1 · 摘自论文原文
  • 引入不合理提示与音画一致性细粒度评估
  • 五款顶尖模型在动作绑定和音画同步上表现不佳
  • 适合研究视频生成局限性与评测方法的学者

逼真文本到视频(T2V)生成技术的快速发展带来了对最新评估方法的迫切需求。现有基准大多忽视了不合理场景,且未衡量音画一致性。我们提出BRITE,首个将(1)不合理提示、(2)音画一致性的细粒度评估、(3)基于问答的可解释评估统一起来的综合性T2V评测框架。不同于易产生幻觉和提示模糊的全自动多模态大模型流水线,BRITE通过严格的人工参与流程确保评测可靠性。评估五款领先模型(Sora 2、Veo 3.1、Runway Gen4.5、Pixverse V5.5、Qwen3Max)发现:尽管模型在静态物体组合上表现良好,但在物体-动作绑定和音画同步方面存在显著退化。该框架为社区提供了一个可靠、可解释的评测工具,能够识别并定位下一代T2V模型在非流形提示下的缺陷。

原文摘要 · Abstract (English)

The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing benchmarks largely overlooked implausible scenarios and do not measure audio-visual alignment. We introduce BRITE, the first framework that unifies (1) implausible prompting, (2) fine-grained assessment of audio-visual consistency, and (3) QA-based interpretable evaluation into a comprehensive T2V benchmark. Unlike fully automated Multimodal LLM-based pipelines, which are prone to hallucination and prompt ambiguity, BRITE guarantees reliability through a rigorous human-in-the-loop protocol for benchmark creation. Evaluating five state-of-the-art models (Sora 2, Veo 3.1, Runway Gen4.5, Pixverse V5.5, and Qwen3Max), we reveal a critical performance gap: while models excel at static object composition, they exhibit significant degradation in object-action binding and audio-visual synchronization. Our framework offers the community a reliable, interpretable benchmark and evaluation framework that can detect and locate limitations in the next generation of T2V models, especially for off-manifold prompts

文本生成视频音画对齐评测基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。