构建首个马拉地语句相似度数据集并训练专用模型
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
- 构建16,860对马拉地语句子的相似度标注数据集
- 模型在0-5分范围内实现稳定性能,优于多种基线模型
- 适合低资源语言研究者用于句法相似性任务
我们提出了MahaSTS,一个针对马拉地语的人工标注句子文本相似度(STS)数据集,以及MahaSBERT-STS-v2,一个为回归式相似度评分优化的微调句向量模型。MahaSTS包含16,860对马拉地语句子,相似度标签为0-5的连续分数,并均匀分布在六个分数区间内,以减少标签偏差并提升模型稳定性。我们在该数据集上微调MahaSBERT模型,并与MahaBERT、MuRIL、IndicBERT和IndicSBERT等模型进行对比。实验表明,MahaSTS有效支持马拉地语句相似度任务的训练,凸显了人工标注、针对性微调和结构化监督在低资源场景下的重要性。数据集与模型已公开于https://github.com/l3cube-pune/MarathiNLP。
原文摘要 · Abstract (English)
We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along with MahaSBERT-STS-v2, a fine-tuned Sentence-BERT model optimized for regression-based similarity scoring. The MahaSTS dataset consists of 16,860 Marathi sentence pairs labeled with continuous similarity scores in the range of 0-5. To ensure balanced supervision, the dataset is uniformly distributed across six score-based buckets spanning the full 0-5 range, thus reducing label bias and enhancing model stability. We fine-tune the MahaSBERT model on this dataset and benchmark its performance against other alternatives like MahaBERT, MuRIL, IndicBERT, and IndicSBERT. Our experiments demonstrate that MahaSTS enables effective training for sentence similarity tasks in Marathi, highlighting the impact of human-curated annotations, targeted fine-tuning, and structured supervision in low-resource settings. The dataset and model are publicly shared at https://github.com/l3cube-pune/MarathiNLP
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。