用微调生成模型提升俄语论文关键词选择效果
Exploring Fine-tuned Generative Models for Keyphrase Selection: A Case Study for Russian
- 微调ruT5、ruGPT等生成模型处理俄语论文关键词选择
- mBART在领域内表现最佳,F1提升12.2%,ROUGE-1增9.0%
- 跨领域仍优于基线,适合关注多语言NLP的研究者
关键词选择在学术文本中至关重要,有助于高效信息检索、摘要生成和索引。本文研究如何将微调的生成式Transformer模型应用于俄语科学文本的关键词选择任务。实验对比了ruT5、ruGPT、mT5和mBART四种模型,在数学与计算机科学、历史、医学、语言学四个领域的俄语科研摘要上评估其性能。结果表明,mBART在领域内表现最优,相比三个基线模型,BERTScore提升4.9%,ROUGE-1提升9.0%,F1-score提升12.2%。尽管跨领域性能显著下降,但在部分场景仍超越基线,显示出该方向具有进一步探索与优化的潜力。
原文摘要 · Abstract (English)
Keyphrase selection plays a pivotal role within the domain of scholarly texts, facilitating efficient information retrieval, summarization, and indexing. In this work, we explored how to apply fine-tuned generative transformer-based models to the specific task of keyphrase selection within Russian scientific texts. We experimented with four distinct generative models, such as ruT5, ruGPT, mT5, and mBART, and evaluated their performance in both in-domain and cross-domain settings. The experiments were conducted on the texts of Russian scientific abstracts from four domains: mathematics & computer science, history, medicine, and linguistics. The use of generative models, namely mBART, led to gains in in-domain performance (up to 4.9% in BERTScore, 9.0% in ROUGE-1, and 12.2% in F1-score) over three keyphrase extraction baselines for the Russian language. Although the results for cross-domain usage were significantly lower, they still demonstrated the capability to surpass baseline performances in several cases, underscoring the promising potential for further exploration and refinement in this research field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。