arXiv:2410.08534cs.CVeess.IV2024-10被引 10

针对AI生成内容的视觉质量评估难题,提出新思路与未来方向。

Quality Prediction of AI Generated Images and Videos: Emerging Trends and Opportunities

  • 聚焦生成式AI内容的感知质量评估,突破传统重建质量衡量局限。
  • 指出现有数据集规模小、失真类型单一,难以真实反映生成内容质量。
  • 为研究者提供可落地的评估框架和开放问题,适合图像视频生成领域从业者。

人工智能已深刻影响人类生活,从自动驾驶到智能聊天机器人,再到基于文本生成图像和视频的模型(如text-to-image、image-to-image、image-to-video)。基于AI的图像视频超分辨率、帧插值、去噪和压缩技术已在产业界引发广泛关注,并部分应用于实际产品。然而,要实现广泛普及与接受,生成或增强的内容必须在视觉上准确、符合使用意图并保持高质量,以避免降低用户体验(QoE)。为此,需部署图像质量评估(IQA)和视频质量评估(VQA)模型进行监控与控制。但多数现有IQA/VQA模型仅以原始参考内容为基准衡量“重建”质量,不适用于评估“生成性”伪影。尽管近期有新指标和模型提出,其评估性能受限于数据集过小、代表性不足或失真覆盖不全,且缺乏对“GenAI”质量评估的有效度量标准。本文探讨生成与增强图像视频内容当前的挑战与机遇,重点关注终端用户感知质量,最后讨论开放问题并提出未来研究建议,推动该前沿领域的进展。

原文摘要 · Abstract (English)

The advent of AI has influenced many aspects of human life, from self-driving cars and intelligent chatbots to text-based image and video generation models capable of creating realistic images and videos based on user prompts (text-to-image, image-to-image, and image-to-video). AI-based methods for image and video super resolution, video frame interpolation, denoising, and compression have already gathered significant attention and interest in the industry and some solutions are already being implemented in real-world products and services. However, to achieve widespread integration and acceptance, AI-generated and enhanced content must be visually accurate, adhere to intended use, and maintain high visual quality to avoid degrading the end user's quality of experience (QoE). One way to monitor and control the visual "quality" of AI-generated and -enhanced content is by deploying Image Quality Assessment (IQA) and Video Quality Assessment (VQA) models. However, most existing IQA and VQA models measure visual fidelity in terms of "reconstruction" quality against a pristine reference content and were not designed to assess the quality of "generative" artifacts. To address this, newer metrics and models have recently been proposed, but their performance evaluation and overall efficacy have been limited by datasets that were too small or otherwise lack representative content and/or distortion capacity; and by performance measures that can accurately report the success of an IQA/VQA model for "GenAI". This paper examines the current shortcomings and possibilities presented by AI-generated and enhanced image and video content, with a particular focus on end-user perceived quality. Finally, we discuss open questions and make recommendations for future work on the "GenAI" quality assessment problems, towards further progressing on this interesting and relevant field of research.

质量评估生成式AI感知质量视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。