arXiv:2505.23358cs.CV2025-05

用知识回放提升视觉语言模型生成更具体、有深度的图像描述

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model

  • 结合束搜索与注意力模块,增强图像特征表达和文本多样性
  • 在多个数据集上显著提升知识识别准确率与描述质量
  • 适合需要高精度知识融合的图文生成场景

生成信息丰富且富含知识的图像描述仍是现有图像描述模型面临的主要挑战,这些模型常产生缺乏特异性与上下文深度的通用描述。为解决此问题,我们提出 KRCapVLM——一种基于知识回放的新型图像描述框架,利用视觉语言模型实现。通过引入束搜索解码策略,生成更具多样性和连贯性的描述;在图像编码器中集成注意力机制模块,增强特征表示能力;并采用训练调度器提升训练稳定性,促进更平滑的收敛。上述方法显著提升了描述质量与知识识别能力。实验表明,该模型在知识识别准确性与整体描述质量上均有明显提升,具备更强泛化能力,能有效处理未见知识概念,生成更丰富、更具上下文相关性的描述。结果验证了该方法在多种场景下生成有意义、知识驱动描述的有效性。

原文摘要 · Abstract (English)

Generating informative and knowledge-rich image captions remains a challenge for many existing captioning models, which often produce generic descriptions that lack specificity and contextual depth. To address this limitation, we propose KRCapVLM, a knowledge replay-based novel image captioning framework using vision-language model. We incorporate beam search decoding to generate more diverse and coherent captions. We also integrate attention-based modules into the image encoder to enhance feature representation. Finally, we employ training schedulers to improve stability and ensure smoother convergence during training. These proposals accelerate substantial gains in both caption quality and knowledge recognition. Our proposed model demonstrates clear improvements in both the accuracy of knowledge recognition and the overall quality of generated captions. It shows a stronger ability to generalize to previously unseen knowledge concepts, producing more informative and contextually relevant descriptions. These results indicate the effectiveness of our approach in enhancing the model's capacity to generate meaningful, knowledge-grounded captions across a range of scenarios.

图像描述视觉语言模型知识增强束搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。