arXiv:2410.06428cs.CLcs.AI2024-10被引 1

用机器学习识别德拉维达语混码文本中的压力,效果优于复杂模型。

Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning

  • 直接使用未清洗文本,结合三种文本表示方法进行分类。
  • 在泰米尔语和泰卢固语上分别达到0.734和0.727的宏F1分数。
  • 适合关注低资源语言心理状态检测的研究者参考。

压力是日常常见情绪,但在某些情境下会影响心理健康,因此构建稳健的检测模型至关重要。本研究提出一种系统性方法,用于识别德拉维达语系中的代码混杂文本中的压力状态。研究涵盖两个数据集,分别针对泰米尔语和泰卢固语。该方法强调使用未清洗文本作为基准,以优化未来分类方法,整合多种预处理技术。采用随机森林算法,结合三种文本表示方式:TF-IDF、词级一元语法以及(1+2+3)-元语法字符组合。该方法在两种语言上均表现良好,泰米尔语实现0.734的宏F1分数,泰卢固语为0.727,优于FastText与Transformer等复杂模型。结果表明,未清洗数据对心理状态检测具有价值,同时也揭示了代码混杂文本分类的挑战,提示通过文本清洗、其他预处理或更复杂模型可进一步提升性能。

原文摘要 · Abstract (English)

Stress is a common feeling in daily life, but it can affect mental well-being in some situations, the development of robust detection models is imperative. This study introduces a methodical approach to the stress identification in code-mixed texts for Dravidian languages. The challenge encompassed two datasets, targeting Tamil and Telugu languages respectively. This proposal underscores the importance of using uncleaned text as a benchmark to refine future classification methodologies, incorporating diverse preprocessing techniques. Random Forest algorithm was used, featuring three textual representations: TF-IDF, Uni-grams of words, and a composite of (1+2+3)-Grams of characters. The approach achieved a good performance for both linguistic categories, achieving a Macro F1-score of 0.734 in Tamil and 0.727 in Telugu, overpassing results achieved with different complex techniques such as FastText and Transformer models. The results underscore the value of uncleaned data for mental state detection and the challenges classifying code-mixed texts for stress, indicating the potential for improved performance through cleaning data, other preprocessing techniques, or more complex models.

压力检测代码混杂低资源语言随机森林

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。