arXiv:2412.09084cs.CL2024-12中稿 · COLING 2025被引 11

用图像化方式处理方言,让模型更好理解非标准语言。

Evaluating Pixel Language Models on Non-Standardized Languages

  • 将文本转为图像块,实现连续词汇表征,适合方言生僻词。
  • 在零样本方言评估中,语法和语义任务最高提升26个百分点。
  • 适合研究方言、低资源语言的学者,但不擅长主题分类。

我们探讨了基于像素的模型从标准语言向方言进行迁移学习的潜力。这类模型将文本转化为图像并分割为块,实现连续词汇表示,特别适用于方言数据中常见的未登录词。以德语为例,我们在多种句法和语义任务上对比了基于像素的模型与基于分词的模型。结果显示,在零样本方言评估中,基于像素的模型在词性标注、依存句法分析和意图识别任务上表现更优,某些场景下性能领先达26个百分点,但在标准德语中无显著优势。然而,其在主题分类任务上表现不佳。这些发现凸显了基于像素模型在处理方言数据方面的潜力,但仍需进一步研究其在不同语言情境下的有效性。

原文摘要 · Abstract (English)

We explore the potential of pixel-based models for transfer learning from standard languages to dialects. These models convert text into images that are divided into patches, enabling a continuous vocabulary representation that proves especially useful for out-of-vocabulary words common in dialectal data. Using German as a case study, we compare the performance of pixel-based models to token-based models across various syntactic and semantic tasks. Our results show that pixel-based models outperform token-based models in part-of-speech tagging, dependency parsing and intent detection for zero-shot dialect evaluation by up to 26 percentage points in some scenarios, though not in Standard German. However, pixel-based models fall short in topic classification. These findings emphasize the potential of pixel-based models for handling dialectal data, though further research should be conducted to assess their effectiveness in various linguistic contexts.

方言识别像素建模零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。