arXiv:2510.17405cs.CLcs.AI2025-10被引 1

为20种非洲语言构建首个可扩展的图像描述系统。

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

  • 基于Flickr8k构建多语言数据集,通过上下文感知翻译生成对齐图文。
  • 采用模型集成与动态替换策略,持续保障生成质量。
  • 0.5B参数模型支持低资源非洲语言,推动多模态AI公平性。

多模态AI研究长期集中于高资源语言,阻碍了技术普惠。为此,我们提出AfriCaption,一个面向20种非洲语言的多语言图像描述框架,贡献包括:(i) 基于Flickr8k构建的数据集,通过上下文感知选择与翻译流程生成语义对齐的描述;(ii) 动态、上下文保持的处理管道,利用模型集成与自适应替换机制确保持续高质量输出;(iii) AfriCaption模型,一个0.5B参数的视觉到文本架构,融合SigLIP与NLLB200,在低资源非洲语言中实现图像描述生成。该统一框架保障数据持续质量,建立了首个面向低资源非洲语言的可扩展图像-描述资源,为真正包容的多模态人工智能奠定基础。

原文摘要 · Abstract (English)

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingual image captioning in 20 African languages and our contributions are threefold: (i) a curated dataset built on Flickr8k, featuring semantically aligned captions generated via a context-aware selection and translation process; (ii) a dynamic, context-preserving pipeline that ensures ongoing quality through model ensembling and adaptive substitution; and (iii) the AfriCaption model, a 0.5B parameter vision-to-text architecture that integrates SigLIP and NLLB200 for caption generation across under-represented languages. This unified framework ensures ongoing data quality and establishes the first scalable image-captioning resource for under-represented African languages, laying the groundwork for truly inclusive multimodal AI.

图像描述非洲语言多模态AI低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。