arXiv:2410.00477cs.CV2024-10被引 3

构建视频危险度评估数据集,测试大模型能否像人一样判断危险程度。

ViDAS: Vision-based Danger Assessment and Scoring

  • 收集100段YouTube视频,人工标注每段视频的危险等级和具体危险时刻。
  • 用大模型基于视频摘要评估危险度,与人类标注对比得出平均误差(MSE)。
  • 适合关注视频安全评估、多模态评测和大模型认知能力的研究者。

我们提出一个新数据集,旨在推进视频内容中的危险性分析与评估,解决量化视频危险程度及衡量大语言模型(LLM)评估者是否具有类人能力的挑战。该数据集包含100段来自YouTube的视频,涵盖多种事件。每段视频由人类参与者按0(无危险)到10(危及生命)的量表进行危险评级,并标注出危险程度升高的精确时间戳。同时,我们利用大语言模型对视频摘要进行独立危险度评估。通过引入均方误差(MSE)实现多模态元评估,衡量人类与大模型在危险评估上的一致性。本数据集不仅为视频危险评估提供了新资源,也展示了大模型实现类人评估的潜力。

原文摘要 · Abstract (English)

We present a novel dataset aimed at advancing danger analysis and assessment by addressing the challenge of quantifying danger in video content and identifying how human-like a Large Language Model (LLM) evaluator is for the same. This is achieved by compiling a collection of 100 YouTube videos featuring various events. Each video is annotated by human participants who provided danger ratings on a scale from 0 (no danger to humans) to 10 (life-threatening), with precise timestamps indicating moments of heightened danger. Additionally, we leverage LLMs to independently assess the danger levels in these videos using video summaries. We introduce Mean Squared Error (MSE) scores for multimodal meta-evaluation of the alignment between human and LLM danger assessments. Our dataset not only contributes a new resource for danger assessment in video content but also demonstrates the potential of LLMs in achieving human-like evaluations.

视频评估大模型评测危险识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。