系统梳理AI视频生成评估方法,推动更全面的评价体系发展。
A Survey of AI-Generated Video Evaluation
- 构建AI生成视频评估框架,涵盖画质、语义、意图对齐与真实感
- 指出当前评估存在指标缺失、人工评测成本高、模型自评不足等短板
- 适合研究者与工业界开发者参考,助力视频生成质量提升
AI生成视频能力的快速发展带来了有效评估的重大挑战。与静态图像或文本不同,视频包含复杂的时空动态,需在视频呈现质量、语义信息传递、与人类意图对齐以及虚拟现实与物理世界一致性等方面进行系统性评估。本综述提出新兴领域AI生成视频评估(AIGVE),强调评估应贴近人类感知并准确响应指令。我们系统分析现有评估方法,梳理其优劣与差距,倡导发展更稳健、细致的评估框架,涵盖传统指标评估、人工评测及未来模型主导的评估方式。本综述旨在为学术界与产业界研究人员提供基础知识支持,推动AI生成视频评估方法的持续进步。
原文摘要 · Abstract (English)
The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which may require a more comprehensive and systematic evaluation of its contents in aspects like video presentation quality, semantic information delivery, alignment with human intentions, and the virtual-reality consistency with our physical world. This survey identifies the emerging field of AI-Generated Video Evaluation (AIGVE), highlighting the importance of assessing how well AI-generated videos align with human perception and meet specific instructions. We provide a structured analysis of existing methodologies that could be potentially used to evaluate AI-generated videos. By outlining the strengths and gaps in current approaches, we advocate for the development of more robust and nuanced evaluation frameworks that can handle the complexities of video content, which include not only the conventional metric-based evaluations, but also the current human-involved evaluations, and the future model-centered evaluations. This survey aims to establish a foundational knowledge base for both researchers from academia and practitioners from the industry, facilitating the future advancement of evaluation methods for AI-generated video content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。