arXiv:2509.15883cs.CVcs.AI2025-09

轻量级图像描述模型,通过关系感知提示提升复杂场景理解能力

RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning

  • 从检索文本中挖掘结构化关系语义,识别图像中的异构对象
  • 仅用10.8M参数实现优于以往轻量模型的性能
  • 适合需要高效精准描述复杂图像的应用场景

近期检索增强型图像描述方法引入外部知识以弥补对复杂场景理解的不足。然而现有方法在关系建模上存在两大问题:(1) 语义提示表示粒度过粗,难以捕捉细粒度关系;(2) 缺乏对图像物体及其语义关系的显式建模。为此,我们提出RACap,一种关系感知的检索增强图像描述模型,不仅能从检索文本中挖掘结构化关系语义,还能从图像中识别异构对象。RACap有效检索包含异构视觉信息的结构化关系特征,增强语义一致性和关系表达力。实验表明,该模型仅含10.8M可训练参数,性能优于以往轻量级描述模型。

原文摘要 · Abstract (English)

Recent retrieval-augmented image captioning methods incorporate external knowledge to compensate for the limitations in comprehending complex scenes. However, current approaches face challenges in relation modeling: (1) the representation of semantic prompts is too coarse-grained to capture fine-grained relationships; (2) these methods lack explicit modeling of image objects and their semantic relationships. To address these limitations, we propose RACap, a relation-aware retrieval-augmented model for image captioning, which not only mines structured relation semantics from retrieval captions, but also identifies heterogeneous objects from the image. RACap effectively retrieves structured relation features that contain heterogeneous visual information to enhance the semantic consistency and relational expressiveness. Experimental results show that RACap, with only 10.8M trainable parameters, achieves superior performance compared to previous lightweight captioning models.

图像描述检索增强轻量模型关系建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。