arXiv:2411.16619cs.CV2024-11中稿 · ACMMM 2025被引 19

构建首个针对人类动作生成视频的评估数据集与指标,解决真实场景中视频质量难测问题。

Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metric

  • 构建6000条人类动作生成视频数据集,覆盖15个主流文生视频模型
  • 提出可解释的量化指标GHVQ,显著优于现有方法
  • 适合视频生成、内容安全、人机交互研究者参考

近年来,基于AI的视频生成技术取得了显著进展。然而,涉及人类动作的生成视频(AGVs)常存在显著的视觉与语义失真,限制了其在真实场景中的应用。为此,本文首次开展人类动作生成视频的质量评估研究,聚焦视觉质量与语义失真识别。首先,构建了包含6000条由15个主流文本到视频(T2V)模型生成的视频数据集——Human-AGVQA,基于400个描述多样化人类动作的文本提示生成。通过主观评测,评估了生成视频中人体外观质量、动作连贯性及整体质量,并识别出身体部位的语义错误。基于该数据集,对T2V模型性能进行基准测试,分析其在不同人类动作类别上的优劣势。其次,提出一种客观评价指标GHVQ,系统提取人体关注特征、生成内容感知特征与时间连续性特征,实现对人类动作生成视频的全面且可解释的评估。大量实验表明,GHVQ在Human-AGVQA数据集上显著优于现有指标,验证了其有效性。相关数据集与指标将开源发布于https://github.com/zczhang-sjtu/GHVQ.git。

原文摘要 · Abstract (English)

AI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git.

视频生成质量评估人类动作数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。