用大模型把英文科研数据集转成俄文,省去人工标注
Transferring Natural Language Datasets Between Languages Using Large Language Models for Modern Decision Support and Sci-Tech Analytical Systems
- 用大模型实现跨语言数据与标注迁移
- 在俄语中首次构建含三层标注的术语定义数据集
- 适合做科技趋势分析和多语言NLP研究者
研发决策依赖于特定研究领域的当前趋势信息。本文研究如何利用大语言模型(LLMs)将数据集及其标注从一种语言迁移到另一种语言。此举至关重要,因跨语言知识共享可推动目标语言中资源匮乏的研究方向,大幅节省数据标注成本或加速原型开发。实验以英俄双语对为例,迁移了DEFT(Definition Extraction from Texts)语料库,该语料库包含三层次标注,专用于术语-定义对挖掘,此类标注在俄语中极为罕见。该数据集对科学领域趋势分析的自然语言处理方法具有基础价值,因术语与定义是科学领域的基本单元。本文提供了一套基于LLM的标注迁移流程,并在翻译后的数据集上训练了BERT基线模型。
原文摘要 · Abstract (English)
The decision-making process to rule R&D relies on information related to current trends in particular research areas. In this work, we investigated how one can use large language models (LLMs) to transfer the dataset and its annotation from one language to another. This is crucial since sharing knowledge between different languages could boost certain underresourced directions in the target language, saving lots of effort in data annotation or quick prototyping. We experiment with English and Russian pairs, translating the DEFT (Definition Extraction from Texts) corpus. This corpus contains three layers of annotation dedicated to term-definition pair mining, which is a rare annotation type for Russian. The presence of such a dataset is beneficial for the natural language processing methods of trend analysis in science since the terms and definitions are the basic blocks of any scientific field. We provide a pipeline for the annotation transfer using LLMs. In the end, we train the BERT-based models on the translated dataset to establish a baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。