arXiv:2607.02259cs.CL2026-07

BamiBERT是首个无需分词的越南语BERT模型,性能超越现有主流模型。

BamiBERT: A New BERT-based Language Model for Vietnamese

  • 从头训练129GB语料,支持2048词元长上下文
  • 在8个任务中11项指标领先,多项达新纪录
  • 适合需要高精度越南语理解的应用场景

本文提出BamiBERT,一种全新的越南语BERT预训练语言模型,旨在解决当前主流越南语文本编码器PhoBERT的关键局限。BamiBERT在129GB通用领域越南语语料上从头训练20轮,支持最长2048词元的上下文长度,并可直接处理原始输入,无需外部分词。在8个越南语基准测试中,其在15项指标中有11项取得最佳成绩,3项位居第二,成为“base”规模越南语编码器的新基准,展现出强大的跨领域泛化能力。模型已发布于Hugging Face:https://huggingface.co/Qualcomm-AI-Research/BamiBERT。

原文摘要 · Abstract (English)

In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT -- the current de facto Vietnamese text encoder. Trained from scratch on a 129GB corpus of general-domain Vietnamese text for 20 epochs, BamiBERT supports an extended context length of up to 2048 tokens and operates directly on raw input, eliminating the need for external word segmentation. Across 8 Vietnamese benchmarks, it achieves the best score on 11 of 15 metrics and the second-best on 3 others, setting a new state of the art among "base"-sized Vietnamese encoders and demonstrating strong cross-domain generalization. We release BamiBERT at: https://huggingface.co/Qualcomm-AI-Research/BamiBERT

越南语BERT预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。