对比大模型与微调,发现微调仍更适配土耳其语情感分析。
Do We Still Need Fine Tuning? Turkish Sentiment Analysis in the Era of Large Language Model
- 用微调BERTurk模型比提示大模型在三分类任务中表现更好
- 大模型在二分类中表现尚可,但在三分类时将中性评论误判为正负
- 中性类别的存在对评估模型鲁棒性至关重要
本研究考察在大语言模型时代,土耳其语情感分析是否仍需监督微调。我们在包含负面、中性、正面标签的土耳其电商评论数据集上,对比了传统机器学习方法、微调预训练语言模型和提示大语言模型的表现。微调后的BERTurk模型整体表现最优,在全三分类任务中优于所有提示大语言模型。中性类别成为主要难点:尽管多个大语言模型在二分类(正/负)中表现接近,但在三分类设置下,其性能显著下降,将中性评论大量误归为极化类别。结果表明,在真实的土耳其语情感分类任务中,提示大语言模型尚未在零样本设置下达到微调模型的水平,且包含中性类别的评估对模型鲁棒性判断至关重要。
原文摘要 · Abstract (English)
This study examines whether supervised fine-tuning remains necessary for Turkish sentiment analysis in the era of large language models. We compare classical machine learning methods, fine-tuned pretrained language models, and prompted large language models on a Turkish e-commerce review dataset with negative, neutral, and positive labels. Fine-tuned BERTurk models perform best overall and outperform all prompted large language models in the full three-class task. The neutral class emerges as the main difficulty: while several large language models are much more competitive in binary positive--negative classification, they degrade substantially in the three-class setting by collapsing neutral reviews into polarized categories. The findings suggest that, in realistic Turkish sentiment classification, prompted large language models do not yet match supervised fine-tuning in the zero-shot setting, and that including the neutral class is crucial for robust evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。