构建多标注视频字幕基准,更真实评估模型与人类理解的一致性。
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
- 每段视频由5人独立标注,捕捉人类感知多样性。
- 提出认知加权指标FIOVA-DQ,精准衡量事件重要性与覆盖度。
- 揭示大模型在复杂视频中漏述事件、模板化表达等系统性问题。
尽管大型视觉语言模型(LVLMs)发展迅速,现有视频字幕评测基准仍受限于单次标注和基于词汇相似性的度量,难以反映人类感知的变异性及事件的认知重要性。为此,我们提出FIOVA(Five-In-One Video Annotations)——一个面向人类对齐的评测基准,包含3,002个真实世界视频(平均时长约33.6秒),每个视频由五名标注者独立标注。该设计支持语义多样性与主观一致性建模,为衡量人机对齐提供更丰富基础。我们进一步提出事件级评估指标FIOVA-DQ,融合标注者共识所得的认知权重,实现对事件相关性和语义覆盖度的细粒度评估。基于FIOVA,我们对九种代表性LVLM进行了全面评估,并引入基于标注者间变异性的复杂度分析框架(CV)。结果揭示了不同难度下模型表现的一致性差距,识别出事件描述不足与模板化倾向等结构性缺陷。实验表明FIOVA在诊断长视频字幕模型行为方面具有显著价值,为认知对齐评估树立新标准。数据集、标注、指标与模型输出已公开,支持未来研究。更多信息请访问 https://huuuuusy.github.io/fiova/。
原文摘要 · Abstract (English)
Despite rapid progress in large vision-language models (LVLMs), existing video caption benchmarks remain limited in evaluating their alignment with human understanding. Most rely on a single annotation per video and lexical similarity-based metrics, failing to capture the variability in human perception and the cognitive importance of events. These limitations hinder accurate diagnosis of model capabilities in producing coherent, complete, and human-aligned descriptions. To address this, we introduce FIOVA (Five-In-One Video Annotations), a human-centric benchmark tailored for evaluation. It comprises 3,002 real-world videos (about 33.6s each), each annotated independently by five annotators. This design enables modeling of semantic diversity and inter-subjective agreement, offering a richer foundation for measuring human-machine alignment. We further propose FIOVA-DQ, an event-level evaluation metric that incorporates cognitive weights derived from annotator consensus, providing fine-grained assessment of event relevance and semantic coverage. Leveraging FIOVA, we conduct a comprehensive evaluation of nine representative LVLMs and introduce a complexity-aware analysis framework based on inter-annotator variation (CV). This reveals consistency gaps across difficulty levels and identifies structural issues such as event under-description and template convergence. Our results highlight FIOVA's diagnostic value for understanding LVLM behavior under varying complexity, setting a new standard for cognitively aligned evaluation in long-video captioning. The benchmark, annotations, metric, and model outputs are publicly released to support future evaluation-driven research in video understanding. More detailed information can be found at https://huuuuusy.github.io/fiova/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。