系统梳理跨语言图像描述的注意力模型,指出现有瓶颈与未来方向。
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
- 按架构分类分析注意力机制驱动的跨语言图像描述模型
- 揭示非英语数据稀缺与语义不一致等核心挑战
- 适合关注多语言视觉-语言融合与AI助手研发的研究者
图像描述任务旨在从输入图像生成文本描述,连接计算机视觉与自然语言处理。近年来,基于Transformer的注意力模型通过增强场景理解能力,显著提升了描述生成性能。尽管已有大量深度学习方法的综述,但针对跨语言注意力机制模型的系统性分析仍较匮乏。本文综述了基于注意力机制的图像描述模型,将其分为Transformer-based、deep learning-based和混合型三类;讨论了常用基准数据集及评价指标(如BLEU、METEOR、CIDEr、ROUGE);指出多语言描述中的主要挑战,包括语义不一致、非英语数据稀缺以及推理能力局限。最后提出未来研究方向:多模态学习、面向AI助手、医疗与法医分析的实时应用。本综述为推进注意力驱动的跨语言图像描述研究提供全面参考。
原文摘要 · Abstract (English)
Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption generation by leveraging attention mechanisms for better scene understanding. While various surveys have explored deep learning-based approaches for image captioning, few have comprehensively analyzed attention-based transformer models across multiple languages. This survey reviews attention-based image captioning models, categorizing them into transformer-based, deep learning-based, and hybrid approaches. It explores benchmark datasets, discusses evaluation metrics such as BLEU, METEOR, CIDEr, and ROUGE, and highlights challenges in multilingual captioning. Additionally, this paper identifies key limitations in current models, including semantic inconsistencies, data scarcity in non-English languages, and limitations in reasoning ability. Finally, we outline future research directions, such as multimodal learning, real-time applications in AI-powered assistants, healthcare, and forensic analysis. This survey serves as a comprehensive reference for researchers aiming to advance the field of attention-based image captioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。