arXiv:2505.22273cs.CL2025-05EMNLP被引 1

针对无分词语言,提出高效准确的词汇归一化方法

Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languages

  • 基于预训练模型设计边界感知的归一化方法
  • 在多领域日语数据集上达到高准确率与高效率
  • 首次系统评估多种视角下的归一化性能

词汇归一化研究旨在处理用户生成文本中的非正式表达,但缺乏全面评估,导致难以判断哪些方法在多维度表现优异。针对无分词语言,本文做出三项关键贡献:(1) 构建大规模、多领域的日语归一化数据集;(2) 开发基于先进预训练模型的归一化方法;(3) 在多个评估视角下开展实验。实验表明,编码器仅用与解码器仅用方法在准确率和效率方面均表现良好。

原文摘要 · Abstract (English)

Lexical normalization research has sought to tackle the challenge of processing informal expressions in user-generated text, yet the absence of comprehensive evaluations leaves it unclear which methods excel across multiple perspectives. Focusing on unsegmented languages, we make three key contributions: (1) creating a large-scale, multi-domain Japanese normalization dataset, (2) developing normalization methods based on state-of-the-art pretrained models, and (3) conducting experiments across multiple evaluation perspectives. Our experiments show that both encoder-only and decoder-only approaches achieve promising results in both accuracy and efficiency.

词汇归一化日语处理预训练模型无分词语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。