arXiv:2503.10997cs.CLcs.AI2025-03中稿 · NAACL被引 4

用连贯关系控制生成图像描述,让内容更自然多样。

RONA: Pragmatically Diverse Image Captioning with Coherence Relations

  • 通过连贯关系设计提示,实现语义层面的多样性控制
  • 在多个数据集上优于基线模型,且与真实描述更匹配
  • 适合需要自然、多样化描述的应用场景

写作助手(如Grammarly、Microsoft Copilot)传统上通过句法和语义变化生成多样化的图像描述。然而,人类撰写的描述更注重传达核心信息,并结合语用线索进行表达。为提升描述多样性,有必要探索与视觉内容结合的新型信息传递方式。我们提出RONA,一种针对多模态大语言模型(MLLM)的新颖提示策略,利用连贯关系作为可控制的语用变化轴。实验表明,相较于多个领域的MLLM基线模型,RONA生成的描述具有更高的整体多样性与真实标注对齐度。代码已开源:https://github.com/aashish2000/RONA。

原文摘要 · Abstract (English)

Writing Assistants (e.g., Grammarly, Microsoft Copilot) traditionally generate diverse image captions by employing syntactic and semantic variations to describe image components. However, human-written captions prioritize conveying a central message alongside visual descriptions using pragmatic cues. To enhance caption diversity, it is essential to explore alternative ways of communicating these messages in conjunction with visual content. We propose RONA, a novel prompting strategy for Multi-modal Large Language Models (MLLM) that leverages Coherence Relations as a controllable axis for pragmatic variations. We demonstrate that RONA generates captions with better overall diversity and ground-truth alignment, compared to MLLM baselines across multiple domains. Our code is available at: https://github.com/aashish2000/RONA

图像描述多模态提示工程多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。