arXiv:2602.01246cs.CLcs.IR2026-02被引 1

首个面向波斯语的开放域推理问答基准,助力低资源语言模型评估

PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian

  • 基于可控LLM生成与人工校验构建波斯语推理题集
  • 包含10,800道多类型题目,覆盖布尔、选择与事实型问题
  • 支持多语言模型对比与波斯语专用模型优化,适合低资源语言研究

推理导向的问答(QA)在大语言模型(LLMs)推动下迅速发展,但低资源语言的高质量基准仍严重匮乏。波斯语作为约1.3亿人使用的语言,缺乏全面的开放域推理问答资源。本文提出PARSE,首个面向波斯语的开放域推理问答基准,包含10,800个问题,涵盖布尔、多项选择和事实型三种格式,涵盖多样化的推理类型、难度层级与答案结构。该基准通过受控的LLM生成流程构建,并经由人工评估验证。我们采用多阶段过滤、标注与一致性检查保障语言与事实质量。在多种提示策略下对多语言及波斯语专用模型进行基准测试,结果表明使用波斯语提示与结构化提示(布尔/选择题用思维链,事实题用少样本)可显著提升性能。微调进一步提升结果,尤其对专精波斯语的模型效果更佳。这些发现凸显PARSE在实现公平比较与实际模型适配方面的价值。PARSE填补了波斯语问答研究的关键空白,为低资源环境下推理能力模型的开发与评估提供了坚实基础。

原文摘要 · Abstract (English)

Reasoning-focused Question Answering (QA) has advanced rapidly with Large Language Models (LLMs), yet high-quality benchmarks for low-resource languages remain scarce. Persian, spoken by roughly 130 million people, lacks a comprehensive open-domain resource for evaluating reasoning-capable QA systems. We introduce PARSE, the first open-domain Persian reasoning QA benchmark, containing 10,800 questions across Boolean, multiple-choice, and factoid formats, with diverse reasoning types, difficulty levels, and answer structures. The benchmark is built via a controlled LLM-based generation pipeline and validated through human evaluation. We also ensure linguistic and factual quality through multi-stage filtering, annotation, and consistency checks. We benchmark multilingual and Persian LLMs under multiple prompting strategies and show that Persian prompts and structured prompting (CoT for Boolean/multiple-choice; few-shot for factoid) improve performance. Fine-tuning further boosts results, especially for Persian-specialized models. These findings highlight how PARSE supports both fair comparison and practical model adaptation. PARSE fills a critical gap in Persian QA research and provides a strong foundation for developing and evaluating reasoning-capable LLMs in low-resource settings.

波斯语推理问答低资源语言大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。