新评估框架可全面衡量视频编辑的语义、空间与时间一致性。
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
- 用视觉语言模型提取语义,目标检测追踪主体,大模型精修对象。
- 结合时序一致性和人工评分,实现对编辑质量的多维度量化。
- 适合研究视频生成与编辑的学者,尤其关注真实感与连贯性。
视频编辑模型虽已取得显著进展,但其性能评估仍具挑战。传统指标如CLIP文本与图像分数存在局限:文本分数受限于训练数据不足与层级依赖关系,图像分数则无法评估时序一致性。本文提出SST-EM(语义、空间与时间评估度量),一个基于现代视觉-语言模型(VLM)、目标检测与时序一致性检查的新型评估框架。该框架包含四个组件:(1) 利用VLM从帧中提取语义信息;(2) 通过目标检测进行主要对象追踪;(3) 由大语言模型(LLM)代理实现焦点对象精细化;(4) 基于视觉变压器(ViT)进行时序一致性评估。各组件整合为统一度量,权重由人工评估与回归分析确定。名称SST-EM体现了其对视频评估中语义、空间与时间三方面的聚焦。SST-EM能全面评估视频编辑中的语义保真度与时序流畅性。源代码可在GitHub仓库获取。
原文摘要 · Abstract (English)
Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: text scores are limited by inadequate training data and hierarchical dependencies, while image scores fail to assess temporal consistency. We present SST-EM (Semantic, Spatial, and Temporal Evaluation Metric), a novel evaluation framework that leverages modern Vision-Language Models (VLMs), Object Detection, and Temporal Consistency checks. SST-EM comprises four components: (1) semantic extraction from frames using a VLM, (2) primary object tracking with Object Detection, (3) focused object refinement via an LLM agent, and (4) temporal consistency assessment using a Vision Transformer (ViT). These components are integrated into a unified metric with weights derived from human evaluations and regression analysis. The name SST-EM reflects its focus on Semantic, Spatial, and Temporal aspects of video evaluation. SST-EM provides a comprehensive evaluation of semantic fidelity and temporal smoothness in video editing. The source code is available in the \textbf{\href{https://github.com/custommetrics-sst/SST_CustomEvaluationMetrics.git}{GitHub Repository}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。