对比多种方法提升斯洛伐克语句子相似度计算效果。
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
- 用传统算法特征训练机器学习模型,优化参数选特征。
- 斯洛伐克语下预训练BERT表现优于通用大模型。
- 适合语言资源少的NLP任务研究者参考。
语义文本相似度(STS)在众多自然语言处理任务中起关键作用。尽管在高资源语言中研究广泛,但对斯洛伐克语等低资源语言仍具挑战性。本文对比评估了应用于斯洛伐克语的多种句级STS方法,包括传统算法、监督机器学习模型及第三方深度学习工具。我们利用传统算法输出作为特征,训练多个机器学习模型,并通过人工蜂群优化算法联合进行特征选择与超参数调优。最后评估了若干第三方工具,包括CloudNLP微调模型、OpenAI嵌入模型、GPT-4及预训练的SlovakBERT模型。结果揭示了不同方法间的权衡关系。
原文摘要 · Abstract (English)
Semantic textual similarity (STS) plays a crucial role in many natural language processing tasks. While extensively studied in high-resource languages, STS remains challenging for under-resourced languages such as Slovak. This paper presents a comparative evaluation of sentence-level STS methods applied to Slovak, including traditional algorithms, supervised machine learning models, and third-party deep learning tools. We trained several machine learning models using outputs from traditional algorithms as features, with feature selection and hyperparameter tuning jointly guided by artificial bee colony optimization. Finally, we evaluated several third-party tools, including fine-tuned model by CloudNLP, OpenAI's embedding models, GPT-4 model, and pretrained SlovakBERT model. Our findings highlight the trade-offs between different approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。