针对Sora生成的复杂视频,提出新型评估方法CRAVE。
Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency
- 多粒度文本-时间融合对齐长文本语义与视频动态
- 混合运动保真建模有效检测时序伪影
- 专为下一代AIGC设计,适合研究视频生成质量
以Sora为代表的新一代视频生成模型显著减少了闪烁伪影,支持更长、更复杂的文本提示,并生成具有丰富多样运动模式的长视频。传统VQA方法因针对简单文本和基础运动模式,难以评估此类内容丰富的视频。为此,我们提出面向Sora时代的AIGC视频评估框架CRAVE,通过多粒度文本-时间融合对齐长文本语义与视频动态,并采用混合运动保真建模评估时序伪影。此外,针对现有数据集提示简单、内容单一的问题,我们构建了CRAVE-DB基准数据集,包含新一代模型生成的高内容密度视频及复杂提示。大量实验表明,CRAVE在多个AIGC VQA基准上表现优异,与人类感知高度一致。代码与数据将公开于https://github.com/littlespray/CRAVE。
原文摘要 · Abstract (English)
The advent of next-generation video generation models like \textit{Sora} poses challenges for AI-generated content (AIGC) video quality assessment (VQA). These models substantially mitigate flickering artifacts prevalent in prior models, enable longer and complex text prompts and generate longer videos with intricate, diverse motion patterns. Conventional VQA methods designed for simple text and basic motion patterns struggle to evaluate these content-rich videos. To this end, we propose \textbf{CRAVE} (\underline{C}ontent-\underline{R}ich \underline{A}IGC \underline{V}ideo \underline{E}valuator), specifically for the evaluation of Sora-era AIGC videos. CRAVE proposes the multi-granularity text-temporal fusion that aligns long-form complex textual semantics with video dynamics. Additionally, CRAVE leverages the hybrid motion-fidelity modeling to assess temporal artifacts. Furthermore, given the straightforward prompts and content in current AIGC VQA datasets, we introduce \textbf{CRAVE-DB}, a benchmark featuring content-rich videos from next-generation models paired with elaborate prompts. Extensive experiments have shown that the proposed CRAVE achieves excellent results on multiple AIGC VQA benchmarks, demonstrating a high degree of alignment with human perception. All data and code will be publicly available at https://github.com/littlespray/CRAVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。