构建5000个AI术语多语言数据集,提升非英语研究者可读性
Towards Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST)
- 用大模型+人工结合方式提取2000-2023年顶会5000个术语
- 五种语言翻译经众包验证,准确率优于现有资源
- 无需重训即可提升翻译质量,适合跨语言AI协作
机器翻译虽取得显著进展,但特定领域术语翻译(尤其在人工智能领域)仍具挑战。本文提出GIST,一个大规模多语言AI术语数据集,包含从2000至2023年顶级人工智能会议论文中提取的5000个术语。这些术语被翻译成阿拉伯语、中文、法语、日语和俄语,采用结合大语言模型提取与人工专家翻译的混合框架。数据集质量通过众包评估进行基准测试,结果显示其翻译准确性优于现有资源。将GIST集成到翻译流程中,使用无需重新训练的后处理优化方法,大模型提示可稳定提升BLEU和COMET得分。在ACL Anthology平台的网页演示展示了其实际应用价值,提升了非英语研究者的可访问性。本工作旨在填补人工智能术语资源的关键空白,促进全球范围内的包容性与协作。数据集地址:https://huggingface.co/datasets/Jerry999/multilingual-terminology
原文摘要 · Abstract (English)
The field of machine translation has achieved significant advancements, yet domain-specific terminology translation, particularly in AI, remains challenging. We introduce GIST, a large-scale multilingual AI terminology dataset containing 5K terms extracted from top AI conference papers spanning 2000 to 2023. The terms are translated into Arabic, Chinese, French, Japanese, and Russian using a hybrid framework that combines LLMs for extraction with human expertise for translation. The dataset's quality is benchmarked against existing resources, demonstrating superior translation accuracy through crowdsourced evaluation. GIST is integrated into translation workflows using post-translation refinement methods that require no retraining, where LLM prompting consistently improves BLEU and COMET scores. A web demonstration on the ACL Anthology platform highlights its practical application, showcasing improved accessibility for non-English speakers. This work aims to address critical gaps in AI terminology resources and fosters global inclusivity and collaboration in AI research. Our data is at https://huggingface.co/datasets/Jerry999/multilingual-terminology
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。