构建视频早期风险信号识别基准,测试模型预判危险的能力。
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
- 定义最早风险信号片段,模拟真实预警场景。
- 现有模型在早期信号下预测准确率显著偏低。
- 适合关注视频安全预警与智能监控的研究者。
随着以视频为中心的社交媒体迅速发展,从视觉数据中预判潜在风险事件成为保障公共安全、预防现实事故的重要方向。以往研究多聚焦于驾驶、抗议、自然灾害等领域的监督式视频风险评估,但多数数据集允许模型接触完整视频序列,包括事故发生时刻,大幅降低了任务难度。为更贴近真实场景,我们提出新基准 RiskCueBench,对视频进行精细标注,识别出最早能提示安全隐患的“风险信号片段”。实验结果揭示当前系统在解读动态情境并从早期视觉信号预判未来风险方面存在显著差距,凸显了实际部署视频风险预测模型所面临的关键挑战。
原文摘要 · Abstract (English)
With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing real world accidents. Prior work has extensively studied supervised video risk assessment across domains such as driving, protests, and natural disasters. However, many existing datasets provide models with access to the full video sequence, including the accident itself, which substantially reduces the difficulty of the task. To better reflect real world conditions, we introduce a new video understanding benchmark RiskCueBench in which videos are carefully annotated to identify a risk signal clip, defined as the earliest moment that indicates a potential safety concern. Experimental results reveal a significant gap in current systems ability to interpret evolving situations and anticipate future risky events from early visual signals, highlighting important challenges for deploying video risk prediction models in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。