arXiv:2509.20045cs.CLcs.AI2025-09EMNLP被引 12

揭示多语言模型在方言任务中的分词与表征偏差对性能的影响

Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks

  • 以分词均等性和信息均等性衡量模型表征偏差
  • 分词偏差更影响句法形态任务,信息偏差更影响语义任务
  • 揭露大模型语言支持背后的字符级不匹配问题

方言数据具有人类感知差异小但显著影响模型性能的语言变异。尽管以往研究将方言差距归因于数据量、经济社会因素等,但其影响并不一致。本文直接考察更根本的成因:将预训练多语言模型的分词均等性(TP)和信息均等性(IP)作为表征偏差的度量,与下游任务表现相关联。我们在三种任务上对比了先进的解码器仅模型与编码器模型:方言分类、主题分类和抽取式问答,控制拉丁与非拉丁书写系统、资源丰富与稀缺两种条件。分析显示,TP是依赖句法和形态线索任务(如抽取式问答)的更好预测指标,而IP则更适用于语义任务(如主题分类)。补充分析包括分词器行为、词汇覆盖率及定性洞察,揭示大模型宣称的语言支持常掩盖脚本或分词层级的深层不匹配。

原文摘要 · Abstract (English)

Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e.g., data size, economic and social factors) whose impact, however, turns out to be inconsistent. In this work, we investigate factors impacting the model performance more directly: we correlate Tokenization Parity (TP) and Information Parity (IP), as measures of representational biases in pre-trained multilingual models, with the downstream performance. We compare state-of-the-art decoder-only LLMs with encoder-based models across three tasks: dialect classification, topic classification, and extractive question answering, controlling for varying scripts (Latin vs. non-Latin) and resource availability (high vs. low). Our analysis reveals that TP is a better predictor of the performance on tasks reliant on syntactic and morphological cues (e.g., extractive QA), while IP better predicts performance in semantic tasks (e.g., topic classification). Complementary analyses, including tokenizer behavior, vocabulary coverage, and qualitative insights, reveal that the language support claims of LLMs often might mask deeper mismatches at the script or token level.

多语言方言表征偏差分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。