用大模型检测天城文语言、仇恨言论及攻击目标,效果优异。
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
- 融合MuRIL、IndicBERT等大模型,用焦点损失缓解数据不平衡。
- 三项任务F1分别达0.9980、0.7652、0.6804,表现优异。
- 适合关注印地语等天城文语言处理的研究者与实践者。
本文详细描述了我们在CHiPSAL 2025共享任务中的参赛系统,聚焦于天城文脚本语言的语言检测、仇恨言论识别和目标检测。我们尝试了MuRIL、IndicBERT和Gemma-2等大语言模型及其集成方法,并采用焦点损失等独特技术,应对天城文语言在多语言处理和类别不平衡方面的自然语言理解挑战。我们的方法在所有任务中均取得具有竞争力的结果:子任务A、B、C的F1分数分别为0.9980、0.7652和0.6804。该工作揭示了变压器模型在具有领域特定和语言挑战的任务中的有效性,也为未来改进提供了方向。
原文摘要 · Abstract (English)
This paper presents a detailed system description of our entry for the CHiPSAL 2025 shared task, focusing on language detection, hate speech identification, and target detection in Devanagari script languages. We experimented with a combination of large language models and their ensembles, including MuRIL, IndicBERT, and Gemma-2, and leveraged unique techniques like focal loss to address challenges in the natural understanding of Devanagari languages, such as multilingual processing and class imbalance. Our approach achieved competitive results across all tasks: F1 of 0.9980, 0.7652, and 0.6804 for Sub-tasks A, B, and C respectively. This work provides insights into the effectiveness of transformer models in tasks with domain-specific and linguistic challenges, as well as areas for potential improvement in future iterations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。