开源工具解决越南语社交媒体用语规范化难题
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
- 结合预训练模型与弱监督学习,实现非标准词汇精准转换
- 支持查询非标准词标准形式,并自动标准化整段文本
- 适合语言研究者和无技术背景用户快速使用
ViSoLex 是一个开源系统,旨在应对越南语社交媒体文本中词汇规范化的独特挑战。该平台提供两项核心服务:非标准词(NSW)查询和词汇规范化,使用户能够获取非正式语言的标准形式,并对含非标准词的文本进行标准化处理。系统架构融合预训练语言模型与弱监督学习技术,确保在越南语标注数据稀缺的情况下仍能实现高精度、高效能的规范化。本文详述了系统的整体设计、功能实现及其对研究人员和非技术用户的适用性。此外,ViSoLex 提供灵活可定制的框架,可适配多种数据集与研究需求。通过发布源代码,系统致力于推动更稳健的越南语自然语言处理工具发展,并鼓励相关领域进一步研究。未来方向包括拓展至更多语言及增强对复杂非标准语言模式的处理能力。
原文摘要 · Abstract (English)
ViSoLex is an open-source system designed to address the unique challenges of lexical normalization for Vietnamese social media text. The platform provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization, enabling users to retrieve standard forms of informal language and standardize text containing NSWs. ViSoLex's architecture integrates pre-trained language models and weakly supervised learning techniques to ensure accurate and efficient normalization, overcoming the scarcity of labeled data in Vietnamese. This paper details the system's design, functionality, and its applications for researchers and non-technical users. Additionally, ViSoLex offers a flexible, customizable framework that can be adapted to various datasets and research requirements. By publishing the source code, ViSoLex aims to contribute to the development of more robust Vietnamese natural language processing tools and encourage further research in lexical normalization. Future directions include expanding the system's capabilities for additional languages and improving the handling of more complex non-standard linguistic patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。