用深度学习融合发音与拼写信息,提升荷兰语切音准确率。
Assessing Dutch Syllabification Algorithms and Improving Accuracy by Combining Phonetic and Orthographic Information through Deep Learning
- 结合发音与拼写特征构建深度学习模型
- 达到99.65%词级准确率,优于现有最佳方法
- 适合语言处理、语音合成等研究者参考
切音是将单词划分为音节的任务。由于规则复杂且存在大量例外,高精度训练切音算法仍是挑战。尽管过去几十年提出了多种荷兰语切音算法,但尚未有全面的比较评估。近年来,深度学习在自然语言处理中日益流行,但尚无针对荷兰语正字法切音的现代深度学习框架。此外,发音与拼写切音算法通常被分别研究,未尝试融合。本研究旨在:(a) 评估现有荷兰语切音算法性能;(b) 探究融合发音与拼写信息是否能提升切音表现。通过在三个数据集(词典词、外来词、伪词)上应用四种算法(Brandt Corstius、Liang、Trogkanis-Elkan (CRF) 及新提出的深度学习模型),发现数据驱动算法在多数条件下优于基于知识的算法。新开发的深度学习模型相比文献最佳结果提升0.14%,达99.65%词级准确率。分析显示,加入发音信息后性能提升的词汇,其拼写歧义可通过发音信息解决。未来可探索发音信息对其他正字法任务的增益,且该框架可拓展至其他语言。
原文摘要 · Abstract (English)
Syllabification describes the task of dividing words into syllables. Due to many rules and exceptions, training an algorithm to perform syllabification with high accuracy remains a challenge. Throughout the last decades, different algorithms have been put forth for Dutch syllabification, yet a comprehensive comparative assessment has not been done. Additionally, deep learning has gained significant popularity within NLP in recent years, yet no modern deep-learning based framework has been developed for Dutch orthographic syllabification. Finally, phonetic and orthographic syllabification algorithms have been examined separately, but not in combination. The aim of the current research was twofold: (a) to examine the performance of existing Dutch syllabification algorithms, and (b) to investigate whether combining phonetic and orthographic information into a single model can increase syllabification performance. To compare the performance of algorithms, four algorithms (Brandt Corstius, Liang, Trogkanis-Elkan (CRF), and a newly conceived deep-learning model) were applied to three different datasets (dictionary words, loanwords, pseudowords). The algorithms show varying performance across datasets, with the data-driven algorithms outperforming a knowledge-based algorithm in all but one condition. The new deep-learning methods developed led to increased performance compared to the best found in the literature (99.65% word accuracy, a 0.14% improvement). An analysis of the words for which adding phonetic information improved syllabification performance indicates that these were words in which the orthographic ambiguity could be resolved by information on pronunciation. Future research could examine other areas where phonetic information can benefit orthographic processing. In addition, the newly developed deep learning frameworks can be applied to other languages than Dutch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。