用检索引导推理,让图像描述更准更细。
Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

- 通过检索获取视觉语言线索,指导模型改进描述。
- 在COCO-LN500上比GRPO提升8.64%关系推理得分。
- 无需额外标注,适合想提升描述质量的开发者。
强化学习在图像描述任务中表现优异,但仍难以推动大视觉语言模型探索新推理策略,导致其性能落后于监督微调。本文提出检索引导的描述优化方法Re$^3$Cap,利用多模态检索作为推理信号,通过描述建议器(CRS)和质量评估器(CQA)识别描述中的幻觉与遗漏,生成更准确、更详细的描述。实验表明,该方法在图像描述任务中优于监督微调,尤其在COCO-LN500基准上相较GRPO平均提升8.64%的关系推理能力。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads to a performance gap between RL and Supervised Fine-Tuning (SFT). In this paper, we argue that multi-modal retrieval can serve as an effective reasoning signal for caption refinement. Based on this insight, we present the Retrieval-Guided Refinement for Image Captioning (Re$^3$Cap), a retrieval-guided reasoning strategy that enhances image captioning without requiring additional annotations. Instantiated by Caption Refinement Suggester (CRS) and Caption Quality Assessor (CQA), this strategy identifies hallucinations and omissions in image captions, leading to more accurate and detailed descriptions. Extensive experiments demonstrate the superiority of our method in image captioning, even compared with Supervised Fine-Tuning. Especially, Re$^3$Cap outperforms GRPO with an average improvement of 8.64% in relation reasoning on the COCO-LN500 benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。