arXiv:2606.19100cs.CV2026-06

首个专为葡萄牙语设计的开源多模态模型,解决语言资源缺失问题。

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

论文配图:AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
图 1 · 摘自论文原文
  • 专为欧洲葡萄牙语构建,使用动态图像分块与定制语言模型
  • 三阶段训练提升多模态对齐与指令理解能力,支持本地化任务
  • 面向研究者与开发者,助力欧洲葡语AI生态建设

大型视觉语言模型发展迅速,但欧洲葡萄牙语(pt-PT)在现有开源多模态模型中长期被忽视,常被误当作巴西葡萄牙语或训练数据严重不足。我们提出AMALIA-VL,首个原生针对pt-PT的开源指令微调多模态模型,结合高分辨率视觉编码器与动态图像分块技术,并通过学习连接器整合完全优化的pt-PT语言模型。我们设计了三阶段训练流程:视觉语言对齐、通用视觉指令微调与偏好优化,并构建以pt-PT为中心的多模态数据集,融合精心筛选与翻译的公开数据及全新采集的本地化数据,填补欧洲葡萄牙语多模态资源几乎空白的现状。评估表明,AMALIA-VL为开源pt-PT多模态模型建立了有力基准。我们将公开模型权重、训练数据及构建流程,并提供机器翻译的pt-PT评估基准,推动欧洲葡萄牙语多模态技术的普及与发展。

原文摘要 · Abstract (English)

Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which either conflate it with Brazilian Portuguese or severely under-represent it in their training data mixes. We introduce AMALIA-VL, the first open-source instruction-tuned LVLM built natively for pt-PT, pairing a high-resolution vision encoder with dynamic image tiling and a fully open pt-PT-optimized language model via a learned connector. We contribute with a purposefully designed three-stage training process - vision-language alignment, general visual instruction tuning, and preference optimization - together with a pt-PT-centric multimodal data mix combining curated and translated public datasets with novel datasets that address the near-total absence of European Portuguese multimodal resources. Our evaluation shows that AMALIA-VL establishes a strong baseline for open-source pt-PT LVLMs. We will release model weights, training data, and construction pipelines along with machine-translated pt-PT evaluation benchmarks to help democratize pt-PT LVLM development.

多模态模型葡萄牙语开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。