arXiv:2510.24813cs.CVcs.AI2025-10

用相似图像生成视觉提示,提升轻量级图像描述的细节捕捉能力

DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts

  • 双检索机制:图文检索+图像间相似场景检索
  • 融合检索图像中的关键物体特征,增强原图视觉表示
  • 参数少但效果优,适合资源受限场景使用

当前轻量级检索增强型图像描述模型通常仅将检索数据用作文本提示,导致原始视觉特征未得到增强,尤其在物体细节或复杂场景下表现不佳。为此,我们提出DualCap,通过从检索到的相似图像中生成视觉提示来丰富视觉表征。模型采用双检索机制:标准图像-文本检索用于生成文本提示,新颖的图像-图像检索则获取视觉上相似的场景。具体而言,从相似场景的描述中提取显著关键词和短语,以捕捉关键物体与相似细节,并将其编码后通过轻量级可训练融合网络与原始图像特征结合。大量实验表明,该方法在参数更少的情况下仍达到优异性能,优于以往视觉提示类描述模型。

原文摘要 · Abstract (English)

Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object details or complex scenes. To address this limitation, we propose $DualCap$, a novel approach that enriches the visual representation by generating a visual prompt from retrieved similar images. Our model employs a dual retrieval mechanism, using standard image-to-text retrieval for text prompts and a novel image-to-image retrieval to source visually analogous scenes. Specifically, salient keywords and phrases are derived from the captions of visually similar scenes to capture key objects and similar details. These textual features are then encoded and integrated with the original image features through a lightweight, trainable feature fusion network. Extensive experiments demonstrate that our method achieves competitive performance while requiring fewer trainable parameters compared to previous visual-prompting captioning approaches.

图像描述视觉提示轻量模型双检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。