用少量人工翻译的波斯语数据,比用大模型生成数据更有效提升低资源语言论点挖掘效果。
Winning with Less for Low Resource Languages: Advantage of Cross-Lingual English_Persian Argument Mining Model over LLM Augmentation
- 结合英波双语数据训练跨语言模型,利用少量人工翻译提升波斯语性能。
- 跨语言模型在波斯语上达74.8%准确率,优于大模型生成数据的69.3%。
- 适合缺乏标注数据但有少量双语资源的语言研究者使用。
论点挖掘是自然语言处理的一个子领域,旨在识别文本中的论点成分(如前提和结论)及其相互关系,揭示文本逻辑结构,用于知识抽取等任务。本文针对低资源语言提出一种跨语言论点挖掘方法,构建三种训练场景:(i) 零样本迁移,仅用英语数据训练;(ii) 英语数据结合大语言模型生成的合成数据增强;(iii) 原始英语数据与人工翻译的波斯语句子联合训练。在英语Microtext语料库及其波斯语平行译本上评估。零样本模型在英语测试集上得50.2% F1,波斯语上50.7%;基于LLM增强的模型分别提升至59.2%和69.3%;而跨语言模型在仅波斯语测试集上达到74.8%的F1,显著优于后者。结果表明,轻量级跨语言融合可大幅超越依赖大模型的数据增强方案,为低资源语言论点挖掘提供可行路径。
原文摘要 · Abstract (English)
Argument mining is a subfield of natural language processing to identify and extract the argument components, like premises and conclusions, within a text and to recognize the relations between them. It reveals the logical structure of texts to be used in tasks like knowledge extraction. This paper aims at utilizing a cross-lingual approach to argument mining for low-resource languages, by constructing three training scenarios. We examine the models on English, as a high-resource language, and Persian, as a low-resource language. To this end, we evaluate the models based on the English Microtext corpus \citep{PeldszusStede2015}, and its parallel Persian translation. The learning scenarios are as follow: (i) zero-shot transfer, where the model is trained solely with the English data, (ii) English-only training enhanced by synthetic examples generated by Large Language Models (LLMs), and (iii) a cross-lingual model that combines the original English data with manually translated Persian sentences. The zero-shot transfer model attains F1 scores of 50.2\% on the English test set and 50.7\% on the Persian test set. LLM-based augmentation model improves the performance up to 59.2\% on English and 69.3\% on Persian. The cross-lingual model, trained on both languages but evaluated solely on the Persian test set, surpasses the LLM-based variant, by achieving a F1 of 74.8\%. Results indicate that a lightweight cross-lingual blend can outperform considerably the more resource-intensive augmentation pipelines, and it offers a practical pathway for the argument mining task to overcome data resource shortage on low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。