arXiv:2508.18904cs.CV2025-08被引 3

构建首个事件级多模态理解大规模基准,推动图像理解从识别到叙事的跃迁。

Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025

  • 基于OpenEvents V1数据集,融合时空语义信息实现事件全要素解析
  • 45支国际团队参与,分两赛道评测,公开与私有测试确保结果可复现
  • 为新闻、媒体分析等场景提供上下文感知的智能图像理解新范式

在ACM Multimedia 2025举办的事件增强图像分析(EVENTA)挑战赛,首次推出面向事件级多模态理解的大规模基准。传统图文生成与检索任务多聚焦于人物、物体和场景的表层识别,常忽视真实事件中的上下文与语义维度。EVENTA通过整合上下文、时间与语义信息,捕捉图像背后的“谁、何时、何地、何事、为何”。基于OpenEvents V1数据集,挑战赛设两个赛道:事件增强图像检索与生成,以及基于事件的图像检索。共有来自六个国家的45支团队参与,通过公开与私有测试阶段评估,确保公平性与可复现性。前三名团队受邀在大会上展示解决方案。EVENTA为上下文感知、叙事驱动的多媒体AI奠定基础,适用于新闻、媒体分析、文化存档及无障碍应用。更多详情见官网:https://ltnghia.github.io/eventa/eventa-2025。

原文摘要 · Abstract (English)

The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses this gap by integrating contextual, temporal, and semantic information to capture the who, when, where, what, and why behind an image. Built upon the OpenEvents V1 dataset, the challenge features two tracks: Event-Enriched Image Retrieval and Captioning, and Event-Based Image Retrieval. A total of 45 teams from six countries participated, with evaluation conducted through Public and Private Test phases to ensure fairness and reproducibility. The top three teams were invited to present their solutions at ACM Multimedia 2025. EVENTA establishes a foundation for context-aware, narrative-driven multimedia AI, with applications in journalism, media analysis, cultural archiving, and accessibility. Further details about the challenge are available at the official homepage: https://ltnghia.github.io/eventa/eventa-2025.

多模态理解事件检测图像检索叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。