arXiv:2411.04756cs.CL2024-11

融合语义与统计特征,提升越南语文本可读性评估准确率

A study of Vietnamese readability assessing through semantic and statistical features

  • 结合语义分析与语法词汇统计特征进行可读性评估
  • 联合模型在维语可读性分类上准确率显著高于单一方法
  • 适合从事越南语自然语言处理与教育技术研究者参考

文本难度评估需综合考量影响理解的各种文本特征,但现有越南语研究仅关注统计特征。本文提出一种融合统计与语义特征的新方法。研究使用三个数据集:越南语可读性数据集(ViRead)、OneStopEnglish 和 RACE(后者已翻译为越南语)。采用 PhoBERT、ViDeBERTa、ViBERTa 等先进语言模型进行语义分析,并结合统计方法提取句法与词汇特征。实验对比了 SVM、随机森林和极端梯度提升等机器学习模型,以准确率与 F1 分数为评价指标。结果表明,联合使用语义与统计特征能显著提升可读性分类准确性。研究强调语义与统计双重考量对越南语文本难度评估的重要性,为未来相关研究奠定基础。

原文摘要 · Abstract (English)

Determining the difficulty of a text involves assessing various textual features that may impact the reader's text comprehension, yet current research in Vietnamese has only focused on statistical features. This paper introduces a new approach that integrates statistical and semantic approaches to assessing text readability. Our research utilized three distinct datasets: the Vietnamese Text Readability Dataset (ViRead), OneStopEnglish, and RACE, with the latter two translated into Vietnamese. Advanced semantic analysis methods were employed for the semantic aspect using state-of-the-art language models such as PhoBERT, ViDeBERTa, and ViBERT. In addition, statistical methods were incorporated to extract syntactic and lexical features of the text. We conducted experiments using various machine learning models, including Support Vector Machine (SVM), Random Forest, and Extra Trees and evaluated their performance using accuracy and F1 score metrics. Our results indicate that a joint approach that combines semantic and statistical features significantly enhances the accuracy of readability classification compared to using each method in isolation. The current study emphasizes the importance of considering both statistical and semantic aspects for a more accurate assessment of text difficulty in Vietnamese. This contribution to the field provides insights into the adaptability of advanced language models in the context of Vietnamese text readability. It lays the groundwork for future research in this area.

可读性评估越南语语义分析机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。