arXiv:2501.04513cs.CVcs.CL2025-01被引 2

让模型在生成时模仿人类改写提示,提升低质量图像描述效果。

Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time

  • 用人类改写数据训练推理阶段反馈模型,无需重训原模型。
  • 在德语图像描述和英文风格迁移上达当前最佳表现。
  • 特别适合改进低质量或非英语的图像描述任务。

将自动预测的人类反馈引入生成模型训练已受到广泛关注,但推理阶段的反馈仍较少被研究。典型的训练阶段反馈(如两样本间的选择偏好)难以自然迁移到推理阶段。本文提出一种新型反馈——标题改写,并基于人类标注训练模型以模仿此类反馈。该方法无需重新训练图像描述模型,显著降低计算开销。我们构建了包含纠错型人类改写的数据集,实验表明将该改写模型引入现有图像描述模型的推理过程,可有效提升生成结果,尤其对原始质量较低的描述改善明显。该方法还成功应用于非英语图像描述任务,取得显著提升;在德语图像描述和英文风格迁移上均达到当前最优性能,且通过详尽的人工对比验证揭示了具体改进维度。

原文摘要 · Abstract (English)

Incorporating automatically predicted human feedback into the process of training generative models has attracted substantial recent interest, while feedback at inference time has received less attention. The typical feedback at training time, i.e., preferences of choice given two samples, does not naturally transfer to the inference phase. We introduce a novel type of feedback -- caption reformulations -- and train models to mimic reformulation feedback based on human annotations. Our method does not require training the image captioning model itself, thereby demanding substantially less computational effort. We experiment with two types of reformulation feedback: first, we collect a dataset of human reformulations that correct errors in the generated captions. We find that incorporating reformulation models trained on this data into the inference phase of existing image captioning models results in improved captions, especially when the original captions are of low quality. We apply our method to non-English image captioning, a domain where robust models are less prevalent, and gain substantial improvement. Second, we apply reformulations to style transfer. Quantitative evaluations reveal state-of-the-art performance on German image captioning and English style transfer, while human validation with a detailed comparative framework exposes the specific axes of improvement.

图像描述推理优化风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。