arXiv:2504.10679cs.CLcs.AI2025-04被引 5

针对斯里兰卡银行领域多语种混合内容,提出高效关键词提取与情感分类方法。

Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content

  • 融合多种模型的混合方法提升多语言文本处理精度
  • 对英语和僧伽罗语混合内容的关键词提取准确率达87.4%以上
  • 适合低资源语言环境下金融舆情监控,可扩展至其他领域

银行品牌声誉依赖于对客户在多语种及代码混合内容中意见的深入分析。传统NLP模型在处理如僧伽罗-英语混合这类低资源语言时,常出现误分类或忽略现象,且难以捕捉领域知识。本研究提出一种混合NLP方法,用于提升银行内容中的关键词提取、数据过滤及基于方面的分类效果。英语关键词提取采用微调后的SpaCy NER、FinBERT-based KeyBERT嵌入、YAKE与EmbedRank相结合的方法,综合准确率达91.2%。僧伽罗语及代码混合内容的关键词提取则使用微调的XLM-RoBERTa模型,并结合领域特定的僧伽罗语金融词汇表,准确率为87.4%。为保障数据质量,通过多个模型进行无关评论过滤,其中BERT-base-uncased在英文上达到85.2%,XLM-RoBERTa在僧伽罗语上达88.1%,优于GPT-4o、SVM和基于关键词的过滤方式。方面分类同样采用类似策略,BERT-base-uncased在英语上达到87.4%,XLM-RoBERTa在僧伽罗语上为85.9%,均超过GPT-4和关键词方法。结果表明,微调的Transformer模型在多语言金融文本分析中显著优于传统方法。该框架为代码混合与低资源语言环境下的品牌声誉监测提供了准确且可扩展的解决方案。

原文摘要 · Abstract (English)

Brand reputation in the banking sector is maintained through insightful analysis of customer opinion on code-mixed and multilingual content. Conventional NLP models misclassify or ignore code-mixed text, when mix with low resource languages such as Sinhala-English and fail to capture domain-specific knowledge. This study introduces a hybrid NLP method to improve keyword extraction, content filtering, and aspect-based classification of banking content. Keyword extraction in English is performed with a hybrid approach comprising a fine-tuned SpaCy NER model, FinBERT-based KeyBERT embeddings, YAKE, and EmbedRank, which results in a combined accuracy of 91.2%. Code-mixed and Sinhala keywords are extracted using a fine-tuned XLM-RoBERTa model integrated with a domain-specific Sinhala financial vocabulary, and it results in an accuracy of 87.4%. To ensure data quality, irrelevant comment filtering was performed using several models, with the BERT-base-uncased model achieving 85.2% for English and XLM-RoBERTa 88.1% for Sinhala, which was better than GPT-4o, SVM, and keyword-based filtering. Aspect classification followed the same pattern, with the BERT-base-uncased model achieving 87.4% for English and XLM-RoBERTa 85.9% for Sinhala, both exceeding GPT-4 and keyword-based approaches. These findings confirm that fine-tuned transformer models outperform traditional methods in multilingual financial text analysis. The present framework offers an accurate and scalable solution for brand reputation monitoring in code-mixed and low-resource banking environments.

多语言处理关键词提取金融NLP代码混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。