修复合成数据中的实体错位,提升零样本图像描述准确性
Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

- 用图像检测到的实体引导重写标题,实现细粒度对齐
- 在多个基准上达到顶尖性能,跨域效果显著提升
- 可插入现有流程,适合需要高质量合成数据的场景
零样本图像描述旨在无需标注图文对的情况下生成图像描述。近期方法利用文本到图像模型从纯文本语料中合成训练数据,但多数关注整体数据质量提升。我们观察到,合成图文错位常具结构性与细粒度特征:图文对可能整体合理,却存在遗漏实体或属性错位,从而降低监督可信度。因此,基于全局相似性的图像重匹配或重生成方法虽提升表面合理性,却无法系统修复实体级错位。为此,我们提出ReCap,一个即插即用框架,将合成数据优化从隐式全局匹配转向显式细粒度对齐。具体地,ReCap通过检测到的图像支撑实体引导标题重写,实现更忠实的合成监督。此外,引入自适应动态加权学习策略,在训练中降低不可靠合成对的权重。作为通用框架,ReCap可集成至现有合成数据流程。大量实验表明,ReCap持续提升图文一致性,并在域内与跨域零样本图像描述基准上取得当前最优表现。
原文摘要 · Abstract (English)
Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from text-only corpora, but most focus on improving overall data quality. In contrast, we observe that synthetic image-text misalignment is often structured and fine-grained: pairs may remain globally plausible while containing missing entities or misgrounded attributes, thereby degrading supervision fidelity. As a result, methods based on global similarity for image rematching or regeneration may improve apparent plausibility, but cannot systematically repair entity-level misalignment. To address this issue, we propose ReCap, a plug-and-play framework that shifts synthetic data refinement from implicit global matching to explicit fine-grained realignment. Specifically, ReCap enforces entity-level correspondence by using detected image-supported entities to guide caption rewriting, yielding more faithful synthetic supervision. In addition, we introduce an adaptive dynamic weighted learning strategy to downweight unreliable synthetic pairs during training. As a general framework, ReCap can be integrated into existing synthetic-data pipelines. Extensive experiments show that ReCap consistently improves image-text consistency and achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。