针对AI生成视频帧间质量不一致问题,提出新评估方法。
Advancing Video Quality Assessment for AIGC
- 用均方误差与交叉熵损失结合,减少生成视频帧间质量波动。
- 在AIGC Video数据集上性能超越当前最优,PLCC提升3.1%。
- 适合关注AI视频生成质量评估的研究者和开发者。
近年来,人工智能生成模型在文本、图像及视频生成等领域取得显著进展。然而,文本到视频生成的质量评估仍处于初级阶段,现有视频质量评估(VQA)框架相较于自然视频评估仍显不足。当前VQA方法主要聚焦于自然视频的整体质量评估,难以有效应对生成视频中帧间质量的显著差异。为此,本文提出一种新型损失函数,将均方误差与交叉熵损失相结合,以缓解帧间质量不一致问题。同时,引入创新的S2CNet技术以保留关键内容,并通过对抗训练增强模型泛化能力。实验结果表明,该方法在AIGC Video数据集上的表现优于现有VQA技术,相比之前最先进方法在PLCC指标上提升3.1%。
原文摘要 · Abstract (English)
In recent years, AI generative models have made remarkable progress across various domains, including text generation, image generation, and video generation. However, assessing the quality of text-to-video generation is still in its infancy, and existing evaluation frameworks fall short when compared to those for natural videos. Current video quality assessment (VQA) methods primarily focus on evaluating the overall quality of natural videos and fail to adequately account for the substantial quality discrepancies between frames in generated videos. To address this issue, we propose a novel loss function that combines mean absolute error with cross-entropy loss to mitigate inter-frame quality inconsistencies. Additionally, we introduce the innovative S2CNet technique to retain critical content, while leveraging adversarial training to enhance the model's generalization capabilities. Experimental results demonstrate that our method outperforms existing VQA techniques on the AIGC Video dataset, surpassing the previous state-of-the-art by 3.1% in terms of PLCC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。