针对土耳其语自动标点与大小写修正,测试了五种规模的BERT模型表现。
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
- 基于BERT设计五种规模模型,适配土耳其语特点
- 模型越大,准确率越高,Base模型精度最优
- 为不同资源条件提供选型参考,适合实际应用部署
本文研究了基于BERT的模型在土耳其语文本中自动标点与大小写修正任务中的有效性,涵盖五种不同规模的模型:Tiny、Mini、Small、Medium和Base。各模型的设计与能力均针对土耳其语的特点进行优化,旨在提升性能的同时降低计算开销。研究系统比较了各模型在精确率、召回率和F1分数上的表现,揭示其在不同应用场景下的适用性。结果表明,随着模型规模增大,文本可读性和修正准确性显著提升,其中Base模型达到最高修正精度。本研究为根据用户需求与计算资源选择合适模型提供了全面指南,并建立了实际应用部署框架,以提升书面土耳其语的质量。
原文摘要 · Abstract (English)
This paper investigates the effectiveness of BERT based models for automated punctuation and capitalization corrections in Turkish texts across five distinct model sizes. The models are designated as Tiny, Mini, Small, Medium, and Base. The design and capabilities of each model are tailored to address the specific challenges of the Turkish language, with a focus on optimizing performance while minimizing computational overhead. The study presents a systematic comparison of the performance metrics precision, recall, and F1 score of each model, offering insights into their applicability in diverse operational contexts. The results demonstrate a significant improvement in text readability and accuracy as model size increases, with the Base model achieving the highest correction precision. This research provides a comprehensive guide for selecting the appropriate model size based on specific user needs and computational resources, establishing a framework for deploying these models in real-world applications to enhance the quality of written Turkish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。