arXiv:2409.09659cs.CL2024-09被引 7

用开源大模型做外语写作母语识别,微调后效果接近商用模型。

Leveraging Open-Source Large Language Models for Native Language Identification

  • 用开源大模型微调,替代传统特征工程方法
  • 微调后在标准数据集上达到与闭源模型相当的准确率
  • 适合关注低成本、可复现语言分析的研究者

母语识别(NLI)是根据第二语言写作文本判断作者母语的任务,应用于司法、营销和二语习得等领域。以往基于手工特征的传统机器学习方法在该任务上优于基于Transformer的模型。近期闭源生成式大模型(如GPT-4)在零样本设置下表现出色,包括开放集分类能力。但闭源模型存在成本高、训练数据不透明等问题。本研究探索开源大模型在NLI中的潜力。结果表明,未经微调的开源模型性能不及闭源模型;但经过标注数据微调后,其表现可与商业模型媲美。

原文摘要 · Abstract (English)

Native Language Identification (NLI) - the task of identifying the native language (L1) of a person based on their writing in the second language (L2) - has applications in forensics, marketing, and second language acquisition. Historically, conventional machine learning approaches that heavily rely on extensive feature engineering have outperformed transformer-based language models on this task. Recently, closed-source generative large language models (LLMs), e.g., GPT-4, have demonstrated remarkable performance on NLI in a zero-shot setting, including promising results in open-set classification. However, closed-source LLMs have many disadvantages, such as high costs and undisclosed nature of training data. This study explores the potential of using open-source LLMs for NLI. Our results indicate that open-source LLMs do not reach the accuracy levels of closed-source LLMs when used out-of-the-box. However, when fine-tuned on labeled training data, open-source LLMs can achieve performance comparable to that of commercial LLMs.

语言识别开源模型大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。