arXiv:2508.06853cs.CVcs.AI2025-08Conference of the …

通过增强视觉显著区域提升图文相关性,生成更准确的图像描述。

AGIC: Attention-Guided Image Captioning to Improve Caption Relevance

  • 在特征空间直接放大关键视觉区域,引导文本生成
  • 在Flickr8k和Flickr30k上超越多个前沿模型,推理更快
  • 结合确定性与概率采样,兼顾流畅性与多样性

尽管图像描述任务取得显著进展,生成准确且描述性强的句子仍是长期挑战。本文提出注意力引导的图像描述(AGIC),通过在特征空间直接放大显著视觉区域来引导描述生成。同时引入混合解码策略,结合确定性与概率采样,平衡生成结果的流畅性与多样性。在Flickr8k和Flickr30k数据集上的大量实验表明,AGIC在多个评价指标上达到或超过现有最优模型表现,且具备更快的推理速度。该方法具有良好的可扩展性与可解释性,为图像描述任务提供了一种高效可靠的解决方案。

原文摘要 · Abstract (English)

Despite significant progress in image captioning, generating accurate and descriptive captions remains a long-standing challenge. In this study, we propose Attention-Guided Image Captioning (AGIC), which amplifies salient visual regions directly in the feature space to guide caption generation. We further introduce a hybrid decoding strategy that combines deterministic and probabilistic sampling to balance fluency and diversity. To evaluate AGIC, we conduct extensive experiments on the Flickr8k and Flickr30k datasets. The results show that AGIC matches or surpasses several state-of-the-art models while achieving faster inference. Moreover, AGIC demonstrates strong performance across multiple evaluation metrics, offering a scalable and interpretable solution for image captioning.

图像描述注意力机制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。