梳理视觉事件与语言生成的关联机制,提出未来研究方向。
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
- 将多帧图像的语言生成任务统一为建模视觉事件与语言关系的通用问题。
- 指出当前模型在捕捉时序视觉-语言交互上的不足。
- 适合关注多模态理解、视频描述与智能叙事的研究者阅读。
近年来,视觉基础自然语言处理领域聚焦于图像或视频内容的描述等真实多模态场景。然而,对不同模态间相互作用的本质与程度关注较少。本文认为,从图像序列或帧生成自然语言的任务,本质上是建模随时间演变的视觉事件与语言特征之间复杂关系的更广泛问题。解决此类任务需模型具备识别和管理这些复杂性的能力。我们分析了五个看似不同的任务,论证它们均属于这一通用多模态问题的实例。随后,综述近年相关建模与评估方法,探讨这些任务共有的挑战。基于此视角,我们提炼关键开放问题,并提出若干未来研究方向。
原文摘要 · Abstract (English)
In recent years, a substantial body of work in visually grounded natural language processing has focused on real-life multimodal scenarios such as describing content depicted in images or videos. However, comparatively less attention has been devoted to study the nature and degree of interaction between the different modalities in these scenarios. In this paper, we argue that any task dealing with natural language generation from sequences of images or frames is an instance of the broader, more general problem of modeling the intricate relationships between visual events unfolding over time and the features of the language used to interpret, describe, or narrate them. Therefore, solving these tasks requires models to be capable of identifying and managing such intricacies. We consider five seemingly different tasks, which we argue are compelling instances of this broader multimodal problem. Subsequently, we survey the modeling and evaluation approaches adopted for these tasks in recent years and examine the common set of challenges these tasks pose. Building on this perspective, we identify key open questions and propose several research directions for future investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。