arXiv:2510.22616cs.CLcs.AI2025-10中稿 · IJCNLP-AACL 2025

首个波斯语常识推理大模型数据集,评测语言理解能力。

PerCoR: Evaluating Commonsense Reasoning in Persian via Multiple-Choice Sentence Completion

  • 用连词分段法生成多样句式补全题,覆盖广泛主题。
  • 人类得分89%,最强模型达92.18%,开源模型仅82.51%。
  • 新方法可提升英文数据集难度,适合多语言推理研究。

我们提出了PerCoR(波斯语常识推理),首个大规模波斯语常识推理基准。PerCoR包含10.6万道从四十多个新闻、文化等网络来源抽取的多项选择句式补全题。我们提出一种基于连词的分段策略,生成结构与主题多样、逻辑连贯的句子补全对。为生成高迷惑性干扰项,提出DRESS-AF(通过嵌入相似度评分与对抗过滤的干扰项排名)方法,不依赖生成模型,从正确答案池中筛选最大混淆度选项。人类在PerCoR上得分为89%,OpenAI-o3表现最佳,达92.18%,紧随其后的是Claude-Sonnet-3.7(91.17%)。最强开源模型DeepSeek-R1得分为82.51%,凸显波斯语常识推理任务的挑战性及当前模型的差距。此外,DRESS-AF方法可迁移至英语HellaSwag基准,在不降低人类解题能力的前提下显著增加难度。数据集已公开于https://huggingface.co/datasets/MCINext/PerCoR。

原文摘要 · Abstract (English)

We introduced PerCoR (Persian Commonsense Reasoning), the first large-scale Persian benchmark for commonsense reasoning. PerCoR contains 106K multiple-choice sentence-completion problems drawn from more than forty news, cultural, and other web sources. We introduce a novel conjunction-based segmentation strategy to generate coherent sentence-completion pairs, enabling broad topical and structural diversity. To create challenging distractors, we propose DRESS-AF (Distractor Ranking via Embedding Similarity Scoring and Adversarial Filtering), a generation-free adversarial filtering method that selects distractors from the pool of gold continuations while maximising model confusion. Human annotators score 89% on PerCoR, while OpenAI-o3 achieves the highest performance at 92.18%, followed closely by Claude-Sonnet-3.7 (91.17%). The strongest open-source model, DeepSeek-R1, reaches 82.51%, underscoring both the dataset's difficulty and the remaining performance gap in Persian commonsense reasoning. We further show that DRESS-AF transfers to the English HellaSwag benchmark, increasing its difficulty without hurting human solvability. The dataset is available at https://huggingface.co/datasets/MCINext/PerCoR.

常识推理波斯语多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。