arXiv:2409.15052cs.CLcs.AI2024-09中稿 · the Ninth Conferen…被引 2

用大模型生成对话提升跨语言图文描述,零训练获高分。

Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning

  • 用指令提示生成图像的多轮对话,融合上下文增强描述。
  • 英印任务达37.90 BLEU,英豪和英孟任务登顶排行榜。
  • 无需训练,适合快速部署到低资源语言场景。

本文介绍团队Brotherhood在WMT 2024英语至低分辨率多模态翻译任务中的系统方案,参与英语-印地语、英语-豪萨语、英语-孟加拉语和英语-马拉雅拉姆语四组语言对。提出一种基于多模态大语言模型(GPT-4o和Claude 3.5 Sonnet)的方法,无需传统训练或微调即可提升跨语言图像描述性能。通过指令微调提示,以图像英文标题为额外上下文,生成丰富的上下文化对话;将这些合成对话翻译为目标语言;最后采用加权提示策略,平衡原始英文标题与翻译后对话,生成目标语言描述。该方法在英印挑战集上取得37.90 BLEU得分,英豪语言对在挑战集与评估榜单分别排名第一和第二。进一步在250张图像子集上实验,分析了不同权重设置下BLEU分数与语义相似性的权衡。

原文摘要 · Abstract (English)

In this paper, we describe our system under the team name Brotherhood for the English-to-Lowres Multi-Modal Translation Task. We participate in the multi-modal translation tasks for English-Hindi, English-Hausa, English-Bengali, and English-Malayalam language pairs. We present a method leveraging multi-modal Large Language Models (LLMs), specifically GPT-4o and Claude 3.5 Sonnet, to enhance cross-lingual image captioning without traditional training or fine-tuning. Our approach utilizes instruction-tuned prompting to generate rich, contextual conversations about cropped images, using their English captions as additional context. These synthetic conversations are then translated into the target languages. Finally, we employ a weighted prompting strategy, balancing the original English caption with the translated conversation to generate captions in the target language. This method achieved competitive results, scoring 37.90 BLEU on the English-Hindi Challenge Set and ranking first and second for English-Hausa on the Challenge and Evaluation Leaderboards, respectively. We conduct additional experiments on a subset of 250 images, exploring the trade-offs between BLEU scores and semantic similarity across various weighting schemes.

跨语言描述大模型应用零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。