arXiv:2606.24058cs.CV2026-06

用真实事件知识提升图像描述的上下文准确性

VisChronos: Revolutionizing Image Captioning Through Real-Life Events

论文配图:VisChronos: Revolutionizing Image Captioning Through Real-Life Events
图 1 · 摘自论文原文
  • 结合大模型与密集描述模型,从单图识别真实事件
  • 生成更详细且符合语境的事件描述,提升描述质量
  • 适合关注事件理解与图像叙事的研究者

本文旨在通过利用现实世界中的历史事件作为知识源,弥合视觉内容与自然语言理解之间的语义鸿沟。我们提出VisChronos框架,结合大型语言模型与密集标注模型,从单张输入图像中自动识别并描述真实事件。该框架可生成详尽、上下文相关的事件描述,显著提升生成字幕的描述质量与语境相关性,克服传统方法在捕捉上下文叙事方面的局限。此外,我们构建了新数据集EventCap(https://zenodo.org/records/14004909),专为增强模型对复杂事件的识别与理解能力而设计。用户研究验证了该方案在生成准确、连贯且聚焦事件的描述方面的有效性,为未来事件中心的图像理解研究铺平道路。

原文摘要 · Abstract (English)

This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world as a source of knowledge for caption generation. We propose VisChronos, a novel framework that utilizes large language models and dense captioning models to identify and describe real-life events from a single input image. Our framework can automatically generate detailed and context-aware event descriptions, enhancing the descriptive quality and contextual relevance of generated captions to address the limitations of traditional methods in capturing contextual narratives. Furthermore, we introduce a new dataset, EventCap (https://zenodo.org/records/14004909), specifically constructed using the proposed framework, designed to enhance the model's ability to identify and understand complex events. The user study demonstrates the efficacy of our solution in generating accurate, coherent, and event-focused descriptions, paving the way for future research in event-centric image understanding.

图像描述事件理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。