用GPT提示词解决低资源德拉威语的词级混合语言识别难题
Prompt Engineering Using GPT for Word-Level Code-Mixed Language Identification in Low-Resource Dravidian Languages
- 设计基于GPT的提示工程方法,识别混合语言中的单词归属
- 卡纳达语模型在多数指标上优于泰米尔语模型,表现更稳定
- 适合研究低资源语言处理与提示工程的学者参考
语言识别(LI)是情感分析、机器翻译和信息检索等自然语言处理任务的基础步骤。在印度等多语言社会中,青年群体在社交媒体上常使用本地语言与英语混合的文本,尤其在词内层面出现混合现象,给LI系统带来挑战。德拉威语广泛分布于南印度,具有丰富的形态结构,但在数字平台中代表性不足,常采用罗马或混合书写方式交流。本文针对共享任务,提出一种基于提示的词级语言识别方法。利用GPT-3.5 Turbo评估大模型对词汇分类的能力。结果显示,卡纳达语模型在多数指标上均优于泰米尔语模型,表现出更高的准确率和可靠性;而泰米尔语模型表现中等,尤其在精确率和召回率方面仍有提升空间。
原文摘要 · Abstract (English)
Language Identification (LI) is crucial for various natural language processing tasks, serving as a foundational step in applications such as sentiment analysis, machine translation, and information retrieval. In multilingual societies like India, particularly among the youth engaging on social media, text often exhibits code-mixing, blending local languages with English at different linguistic levels. This phenomenon presents formidable challenges for LI systems, especially when languages intermingle within single words. Dravidian languages, prevalent in southern India, possess rich morphological structures yet suffer from under-representation in digital platforms, leading to the adoption of Roman or hybrid scripts for communication. This paper introduces a prompt based method for a shared task aimed at addressing word-level LI challenges in Dravidian languages. In this work, we leveraged GPT-3.5 Turbo to understand whether the large language models is able to correctly classify words into correct categories. Our findings show that the Kannada model consistently outperformed the Tamil model across most metrics, indicating a higher accuracy and reliability in identifying and categorizing Kannada language instances. In contrast, the Tamil model showed moderate performance, particularly needing improvement in precision and recall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。