构建爱尔兰语与英语双语评测集,评估大模型在低资源语言中的推理能力。
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
- 基于爱尔兰高考题设计12个主题,支持长文本生成与语言保真度评估
- 最佳模型在爱尔兰语中正确率仅55.8%,低于英语的76.2%,且有效回应不足80%
- 首个面向濒危语言的多模态、文化嵌入式开放问答评测基准
大语言模型虽展现出色的知识与推理能力,但在多语言及低资源场景下的表现仍待深入探索。现有评测常存在文化偏见,仅限文本形式,依赖选择题,且对极低资源语言支持不足。为此,我们提出IRLBench,以英语与爱尔兰语并行呈现,后者被联合国教科文组织列为极度濒危语言。该评测集基于2024年爱尔兰高中毕业考试的12个代表性科目构建,支持细粒度领域分析。通过长文本生成任务与官方评分标准结合,不仅评估答案正确性,还衡量语言忠实度。对主流闭源与开源模型的广泛实验显示,模型在爱尔兰语中的表现显著落后:最佳模型有效回应比例不足80%,正确率仅为55.8%,远低于英语的76.2%。我们已公开IRLBench(https://huggingface.co/datasets/ReliableAI/IRLBench)及配套评估代码库(https://github.com/ReML-AI/IRLBench),推动更具鲁棒性与文化敏感性的多语言AI研究。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have demonstrated promising knowledge and reasoning abilities, yet their performance in multilingual and low-resource settings remains underexplored. Existing benchmarks often exhibit cultural bias, restrict evaluation to text-only, rely on multiple-choice formats, and, more importantly, are limited for extremely low-resource languages. To address these gaps, we introduce IRLBench, presented in parallel English and Irish, which is considered definitely endangered by UNESCO. Our benchmark consists of 12 representative subjects developed from the 2024 Irish Leaving Certificate exams, enabling fine-grained analysis of model capabilities across domains. By framing the task as long-form generation and leveraging the official marking scheme, it does not only support a comprehensive evaluation of correctness but also language fidelity. Our extensive experiments of leading closed-source and open-source LLMs reveal a persistent performance gap between English and Irish, in which models produce valid Irish responses less than 80\% of the time, and answer correctly 55.8\% of the time compared to 76.2\% in English for the best-performing model. We release IRLBench (https://huggingface.co/datasets/ReliableAI/IRLBench) and an accompanying evaluation codebase (https://github.com/ReML-AI/IRLBench) to enable future research on robust, culturally aware multilingual AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。