用视觉相似文本增强生成,零样本图文描述更准
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

- 用图像相似文本检索对齐文本与视觉特征
- 融合检索结果提升描述准确率,视频/图像任务均超越现有方法
- 基于频率过滤实体,适合零样本场景下提升描述质量
近期图像描述研究探索了仅使用文本训练的方法,以克服配对图文数据的局限。然而,现有纯文本训练方法常忽略训练时使用文本与推理时使用图像之间的模态差距。为此,我们提出一种新方法——图像相似文本检索(Image-like Retrieval),将文本特征与视觉相关特征对齐,缓解模态差距。该方法进一步通过设计融合模块,将检索到的描述与输入特征结合,提升生成描述的准确性。此外,我们引入基于频率的实体过滤技术,显著改善描述质量。上述方法整合为统一框架,称为IFCap(Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning)。大量实验表明,该简单而强大的方法在零样本图文描述和视频描述任务中均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To address this issue, we propose a novel approach called Image-like Retrieval, which aligns text features with visually relevant features to mitigate the modality gap. Our method further enhances the accuracy of generated captions by designing a Fusion Module that integrates retrieved captions with input features. Additionally, we introduce a Frequency-based Entity Filtering technique that significantly improves caption quality. We integrate these methods into a unified framework, which we refer to as IFCap ($\textbf{I}$mage-like Retrieval and $\textbf{F}$requency-based Entity Filtering for Zero-shot $\textbf{Cap}$tioning). Through extensive experimentation, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in both image captioning and video captioning compared to zero-shot captioning based on text-only training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。