系统梳理46个葡萄牙语语言模型,揭示研究现状与未来方向
Language Models for Portuguese: A Systematic Mapping Study
- 系统收集并分类46个葡萄牙语语言模型,涵盖架构、数据、许可等维度
- 通过演化关系分析发现模型间存在技术继承与分化趋势
- 为研究者和开发者提供清晰的资源地图,助力后续模型研发
近年来,语言模型的快速发展推动了自然语言处理领域的广泛应用。然而,语言模型的发展在不同语言间并不均衡。以葡萄牙语为例,学术界与企业正积极开发语言模型并构建相关数据资源,形成了日益多元的模型生态。但这些模型的信息分散于论文、技术报告、模型仓库和项目文档中。本综述开展了一项系统映射研究,全面梳理了46个针对葡萄牙语的语言模型,从基础模型、架构、计算资源、训练数据集、许可协议、代码与模型权重可获取性等方面进行特征刻画。进一步通过演化视角分析模型间的关联与发展脉络,识别出当前的研究空白与机遇,并探讨未来发展方向。
原文摘要 · Abstract (English)
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。