构建首个孟加拉语多步推理数据集,评估模型跨语言推理能力。
Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
- 手动翻译英文数据集生成孟加拉语多步推理题。
- 模型在非二元问题上依赖推理步骤更明显,但孟加拉语表现差。
- 适合关注低资源语言与跨语言推理的研究者。
语言模型在复杂多步推理任务中表现出色,但其评估主要集中在英语等高资源语言。本文基于英文Reveal数据集,构建了首个手动翻译的孟加拉语多步推理数据集,包含二元与非二元问题类型。我们在原始数据集和译本上对英語主导与孟加拉語主导的多语言小型模型进行控制实验,比较其利用相关推理步骤得出正确答案的能力。结果显示,在相同条件下,推理上下文对更具挑战性的非二元问题更有帮助,但模型难以有效运用孟加拉语推理步骤。我们进一步分析推理步骤如何影响模型预测,揭示不同模型与语言间的差异趋势。
原文摘要 · Abstract (English)
Language models have demonstrated remarkable performance on complex multi-step reasoning tasks. However, their evaluation has been predominantly confined to high-resource languages such as English. In this paper, we introduce a manually translated Bangla multi-step reasoning dataset derived from the English Reveal dataset, featuring both binary and non-binary question types. We conduct a controlled evaluation of English-centric and Bangla-centric multilingual small language models on the original dataset and our translated version to compare their ability to exploit relevant reasoning steps to produce correct answers. Our results show that, in comparable settings, reasoning context is beneficial for more challenging non-binary questions, but models struggle to employ relevant Bangla reasoning steps effectively. We conclude by exploring how reasoning steps contribute to models' predictions, highlighting different trends across models and languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。