arXiv:2509.25813cs.CLcs.LG2025-09被引 1

构建首个罗马尼亚语生物题数据集,提升大模型在小语种科学领域的理解能力

RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models

  • 构建1.4万道罗马尼亚语生物选择题数据集,专用于评估模型科学理解力
  • 发现当前大模型在小语种科学任务中准确率有限,依赖提示工程可小幅提升
  • 适合关注低资源语言、生物教育、多语种大模型评估的研究者

近年来,大语言模型(LLMs)在自然语言处理任务中展现出显著潜力,但在特定领域和非英语语言中的表现仍待深入探索。本研究推出首个罗马尼亚语生物学选择题数据集,精心设计以评估大模型在科学语境下的理解和推理能力。该数据集包含约14,000个问题,为评估和提升大模型在生物学领域的性能提供了全面资源。我们对多个主流大模型进行了基准测试,分析其准确率、推理模式以及对领域术语和语言细微差别的理解能力。此外,通过系统实验评估了提示工程、微调等优化技术对模型表现的影响。结果揭示了当前大模型在低资源语言科学任务中的优势与局限,为未来研究与发展提供了重要参考。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have demonstrated significant potential across various natural language processing (NLP) tasks. However, their performance in domain-specific applications and non-English languages remains less explored. This study introduces a novel Romanian-language dataset for multiple-choice biology questions, carefully curated to assess LLM comprehension and reasoning capabilities in scientific contexts. Containing approximately 14,000 questions, the dataset provides a comprehensive resource for evaluating and improving LLM performance in biology. We benchmark several popular LLMs, analyzing their accuracy, reasoning patterns, and ability to understand domain-specific terminology and linguistic nuances. Additionally, we perform comprehensive experiments to evaluate the impact of prompt engineering, fine-tuning, and other optimization techniques on model performance. Our findings highlight both the strengths and limitations of current LLMs in handling specialized knowledge tasks in low-resource languages, offering valuable insights for future research and development.

大模型生物题罗马尼亚语低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。