用弱监督框架自动标注越南语社交文本,提升低资源语言的词汇标准化效果。
A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
- 结合半监督与弱监督技术,自动标注非标准词汇为规范形式。
- 在预训练模型下达F1 82.72%,词汇完整性准确率达99.22%。
- 适合处理无变音符号文本,适用于低资源语言NLP任务优化。
本研究提出一种创新的自动标注框架,用于解决越南语等低资源语言在社交媒体文本中的词汇标准化难题。社交媒体数据丰富多样,但其不断演变的语言形式使人工标注成本高昂。为此,我们设计的框架融合半监督学习与弱监督技术,在最小化人工标注的前提下,提升训练数据质量并扩大规模。该框架可自动将原始文本中的非标准词汇转换为标准形式,显著提高训练数据的准确性和一致性。实验表明,该弱监督框架在利用预训练语言模型时表现优异,达到82.72%的F1分数,词汇完整性准确率高达99.22%。同时,对无变音符号文本在多种条件下均有良好处理能力。该框架有效提升了自然语言标准化质量,使各类NLP任务平均准确率提升1%-3%。
原文摘要 · Abstract (English)
This study introduces an innovative automatic labeling framework to address the challenges of lexical normalization in social media texts for low-resource languages like Vietnamese. Social media data is rich and diverse, but the evolving and varied language used in these contexts makes manual labeling labor-intensive and expensive. To tackle these issues, we propose a framework that integrates semi-supervised learning with weak supervision techniques. This approach enhances the quality of training dataset and expands its size while minimizing manual labeling efforts. Our framework automatically labels raw data, converting non-standard vocabulary into standardized forms, thereby improving the accuracy and consistency of the training data. Experimental results demonstrate the effectiveness of our weak supervision framework in normalizing Vietnamese text, especially when utilizing Pre-trained Language Models. The proposed framework achieves an impressive F1-score of 82.72% and maintains vocabulary integrity with an accuracy of up to 99.22%. Additionally, it effectively handles undiacritized text under various conditions. This framework significantly enhances natural language normalization quality and improves the accuracy of various NLP tasks, leading to an average accuracy increase of 1-3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。