arXiv:2409.00358cs.CLcs.AI2024-09NAACL被引 6

用低秩方言适配器提升大模型在方言对话中猜词的能力

Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models

  • 设计低秩方言适配器LoRDD,结合任务与对比学习优化解码器
  • 在印英和尼英对话中,词相似度与准确率差距缩小至5%以下
  • 适合关注方言适应、对话生成或低资源语言处理的研究者

针对特定方言(如印度英语、尼日利亚英语)的自然语言理解任务,已有研究为编码器模型设计方言适配器。本文将该思想扩展至解码器模型,提出名为LoRDD的架构。基于公开的MD-3数据集——包含方言使用者间词语游戏对话——任务为从掩码对话中预测目标词(TWP)。LoRDD融合任务适配器与方言适配器,后者通过在MD-3中构建伪平行对话进行对比学习。在Mistral与Gemma两个模型上,实验表明LoRDD优于四种基线方法。相较于美式英语,其性能差距显著缩小:词相似度差距降至12%和5.8%,准确率差距降至25%和4.5%。本工作首次证明了通过简化的下一个词预测任务(即TWP),实现解码器模型对方言的有效适应。

原文摘要 · Abstract (English)

Dialect adapters that improve the performance of LLMs for NLU tasks on certain sociolects/dialects/national varieties ('dialects' for the sake of brevity) have been reported for encoder models. In this paper, we extend the idea of dialect adapters to decoder models in our architecture called LoRDD. Using MD-3, a publicly available dataset of word game-playing conversations between dialectal speakers, our task is Target Word Prediction (TWP) from a masked conversation. LoRDD combines task adapters and dialect adapters where the latter employ contrastive learning on pseudo-parallel conversations from MD-3. Our experiments on Indian English and Nigerian English conversations with two models (Mistral and Gemma) demonstrate that LoRDD outperforms four baselines on TWP. Additionally, it significantly reduces the performance gap with American English, narrowing it to 12% and 5.8% for word similarity, and 25% and 4.5% for accuracy, respectively. The focused contribution of LoRDD is in its promise for dialect adaptation of decoder models using TWP, a simplified version of the commonly used next-word prediction task.

方言适应解码器模型低秩适配对话预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。