构建首个文本生成视频幻觉评测基准,揭示五大幻觉类型。
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
- 设计五类幻觉分类框架,基于视频嵌入识别生成错误。
- 构建3782条带标注视频数据集,覆盖837个MS COCO图文对。
- 适合研究视频生成可靠性与幻觉检测的学者使用。
大型多模态模型在视频理解方面取得进展,文本到视频(T2V)模型能根据文本提示生成视频,但常产生幻觉内容,导致生成结果不一致。本文提出ViBe:一个大规模开源T2V模型生成的幻觉视频数据集。我们识别出五种主要幻觉类型:主体消失、遗漏错误、数值不一致、主体变形和视觉不符。通过十款T2V模型,从837个多样化的MS COCO图文描述中生成并人工标注了3,782条视频。该基准包含幻觉视频数据集和基于视频嵌入的分类框架。实验以分类任务为基线,TimeSFormer + CNN集成模型表现最佳(准确率0.345,F1分数0.342)。尽管初始基线性能有限,表明自动化幻觉检测难度高,亟需更优方法。本研究旨在推动更鲁棒的T2V模型发展,并基于用户偏好评估输出质量。
原文摘要 · Abstract (English)
Recent advances in Large Multimodal Models (LMMs) have expanded their capabilities to video understanding, with Text-to-Video (T2V) models excelling in generating videos from textual prompts. However, they still frequently produce hallucinated content, revealing AI-generated inconsistencies. We introduce ViBe (https://vibe-t2v-bench.github.io/): a large-scale dataset of hallucinated videos from open-source T2V models. We identify five major hallucination types: Vanishing Subject, Omission Error, Numeric Variability, Subject Dysmorphia, and Visual Incongruity. Using ten T2V models, we generated and manually annotated 3,782 videos from 837 diverse MS COCO captions. Our proposed benchmark includes a dataset of hallucinated videos and a classification framework using video embeddings. ViBe serves as a critical resource for evaluating T2V reliability and advancing hallucination detection. We establish classification as a baseline, with the TimeSFormer + CNN ensemble achieving the best performance (0.345 accuracy, 0.342 F1 score). While initial baselines proposed achieve modest accuracy, this highlights the difficulty of automated hallucination detection and the need for improved methods. Our research aims to drive the development of more robust T2V models and evaluate their outputs based on user preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。