arXiv:2509.09131cs.CLcs.AI2025-09

专为越南语设计的高效重排序模型,提升低资源语言检索效果。

ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking

  • 基于BGE-M3与块并行变压器架构,适配越南语复杂语法。
  • 在MMARCO-VI上超越多语言基线,接近PhoRanker表现。
  • 开源模型支持可复现性,适用于其他低资源语言研究。

本文提出ViRanker,一种针对越南语的交叉编码器重排序模型。基于BGE-M3编码器并引入块并行变压器结构,解决越南语这一低资源语言在语法复杂性和变音符号方面的挑战。模型在8 GB精心筛选的语料库上训练,并采用混合难例采样进行微调,增强鲁棒性。在MMARCO-VI基准测试中,ViRanker展现出优异的早期排名准确率,超越多语言基线,与PhoRanker表现接近。通过在Hugging Face公开发布模型,旨在促进可复现性并推动其在实际检索系统中的应用。本研究还表明,精心的架构适配与数据整理可有效推进其他未充分代表语言的重排序性能。

原文摘要 · Abstract (English)

This paper presents ViRanker, a cross-encoder reranking model tailored to the Vietnamese language. Built on the BGE-M3 encoder and enhanced with the Blockwise Parallel Transformer, ViRanker addresses the lack of competitive rerankers for Vietnamese, a low-resource language with complex syntax and diacritics. The model was trained on an 8 GB curated corpus and fine-tuned with hybrid hard-negative sampling to strengthen robustness. Evaluated on the MMARCO-VI benchmark, ViRanker achieves strong early-rank accuracy, surpassing multilingual baselines and competing closely with PhoRanker. By releasing the model openly on Hugging Face, we aim to support reproducibility and encourage wider adoption in real-world retrieval systems. Beyond Vietnamese, this study illustrates how careful architectural adaptation and data curation can advance reranking in other underrepresented languages.

越南语重排序低资源语言交叉编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。