用检索增强生成文化图像描述,多语言翻译性能显著提升。
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
- 两阶段流程:先生成西班牙语中间描述,再通过检索增强提示生成目标语言描述。
- 在布里布里、瓜拉尼和奥里萨巴纳瓦特语上相比基线提升超120%,最高达164.1%。
- 适合低资源原住民语言的图像描述生成,尤其对有领域语料的语言效果更佳。
我们提交了佛罗里达大学Gators团队参加AmericasNLP 2026关于原住民语言文化图像描述的共享任务。采用两阶段流水线:首先用Qwen2.5-VL生成西班牙语中间描述,再利用Gemini 2.5 Flash进行检索增强的多示例提示生成目标语言描述。在开发集上,布里布里、瓜拉尼和奥里萨巴纳瓦特语的性能分别较基线提升164.1%、131.7%和122.6%;测试集上布里布里和奥里萨巴纳瓦特语仍保持>150%的提升。我们发现检索效果高度依赖语言,仅在大且领域相关的语料中有效,而合成数据增强贡献了约28 chrF++的瓜拉尼语言性能提升。该方案为共享任务总体优胜者,在人类评估中位列五项决赛方案中的第二名。
原文摘要 · Abstract (English)
We present the University of Florida Gators submission to the AmericasNLP 2026 shared task on cultural image captioning for Indigenous languages. Our two-stage pipeline generates a Spanish intermediate caption with Qwen2.5-VL, then produces the target-language caption using retrieval-augmented many-shot prompting with Gemini 2.5 Flash. We achieve 164.1%, 131.7%, and 122.6% improvements over the shared task baseline for Bribri, Guaraní, and Orizaba Nahuatl captioning, respectively, in our dev set evaluation and maintain >150% improvements for the Bribri and Orizaba Nahuatl languages in the test set evaluation. We find retrieval is highly language-dependent, beneficial only for large, in-domain corpora, and that synthetic data augmentation accounts for around 28 chrF++ of the dev set Guaraní performance gain. Our submission is the overall winner of the shared task, placing second out of five finalist submissions in human evaluations of target-language captions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。