针对无分词语言,提出高效准确的词汇归一化方法
Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languages
- 基于预训练模型设计边界感知的归一化方法
- 在多领域日语数据集上达到高准确率与高效率
- 首次系统评估多种视角下的归一化性能
词汇归一化研究旨在处理用户生成文本中的非正式表达,但缺乏全面评估,导致难以判断哪些方法在多维度表现优异。针对无分词语言,本文做出三项关键贡献:(1) 构建大规模、多领域的日语归一化数据集;(2) 开发基于先进预训练模型的归一化方法;(3) 在多个评估视角下开展实验。实验表明,编码器仅用与解码器仅用方法在准确率和效率方面均表现良好。
原文摘要 · Abstract (English)
Lexical normalization research has sought to tackle the challenge of processing informal expressions in user-generated text, yet the absence of comprehensive evaluations leaves it unclear which methods excel across multiple perspectives. Focusing on unsegmented languages, we make three key contributions: (1) creating a large-scale, multi-domain Japanese normalization dataset, (2) developing normalization methods based on state-of-the-art pretrained models, and (3) conducting experiments across multiple evaluation perspectives. Our experiments show that both encoder-only and decoder-only approaches achieve promising results in both accuracy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。