提出新评测基准ARGUS,精准衡量视频大模型幻觉与遗漏问题
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
- 构建自由文本生成评测框架,对比模型与人工标注的视频描述
- 量化两类缺陷:虚构内容占比、重要细节遗漏率
- 适合评估视频大模型真实表现,尤其关注生成可靠性
视频大语言模型尚未广泛部署,主要因其易产生幻觉。现有评测多依赖选择题,但模型在自由文本生成任务(如视频字幕)中幻觉更严重。为此,我们提出ARGUS基准,通过对比模型输出与人工标注字幕,量化双重指标:一是关于视频内容或时序关系的错误陈述比例;二是重要描述信息的遗漏率。二者共同构成对视频字幕生成性能的全面评估。
原文摘要 · Abstract (English)
Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinate far more aggressively on freeform text generation tasks like video captioning than they do on multiple choice verification tasks. To address this weakness, we propose ARGUS, a VideoLLM benchmark that measures freeform video captioning performance. By comparing VideoLLM outputs to human ground truth captions, ARGUS quantifies dual metrics. First, we measure the rate of hallucinations in the form of incorrect statements about video content or temporal relationships. Second, we measure the rate at which the model omits important descriptive details. Together, these dual metrics form a comprehensive view of video captioning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。