构建大规模多模态事件理解数据集,支持事件上下文与时间定位
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- 基于新闻与图片构建事件感知的多模态任务框架
- 涵盖20万+新闻与40万张图片,覆盖多元领域与时间跨度
- 适合研究跨模态推理、事件理解与视觉语言模型的开发者
我们提出 OpenEvents V1,一个大规模基准数据集,旨在推动以事件为中心的视觉-语言理解。与传统聚焦表面描述的图像字幕和检索数据集不同,OpenEvents V1 通过三项核心任务强调上下文与时间定位:(1) 生成富含事件信息的图像字幕,(2) 根据图像查询检索相关新闻文章,(3) 根据叙事性文本查询检索相关图像。数据集包含来自 CNN 与 The Guardian 的超过 20 万篇新闻文章和 40 万张关联图像,覆盖广泛领域与时间范围。我们提供了各项任务的详尽基线结果与标准化评估协议。OpenEvents V1 为开发能够对复杂现实事件进行深度推理的多模态 AI 系统奠定了坚实基础。数据集公开获取地址:https://ltnghia.github.io/eventa/openevents-v1。
原文摘要 · Abstract (English)
We introduce OpenEvents V1a large-scale benchmark dataset designed to advance event-centric vision-language understanding. Unlike conventional image captioning and retrieval datasets that focus on surface-level descriptions, OpenEvents V1 dataset emphasizes contextual and temporal grounding through three primary tasks: (1) generating rich, event-aware image captions, (2) retrieving event-relevant news articles from image queries, and (3) retrieving event-relevant images from narrative-style textual queries. The dataset comprises over 200,000 news articles and 400,000 associated images sourced from CNN and The Guardian, spanning diverse domains and time periods. We provide extensive baseline results and standardized evaluation protocols for all tasks. OpenEvents V1 establishes a robust foundation for developing multimodal AI systems capable of deep reasoning over complex real-world events. The dataset is publicly available at https://ltnghia.github.io/eventa/openevents-v1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。