arXiv:2602.17623cs.CL2026-02被引 1

揭示波斯语大模型在文化常识推理上的认知鸿沟

Unmasking the Factual-Conceptual Gap in Persian Language Models

  • 构建新基准DivanBench,测试迷信与习俗的隐含规则理解
  • 7个模型在情境推理上比事实检索低21%,普遍存在盲从偏差
  • 持续预训练反而加剧偏见,暴露模型未真正内化文化逻辑

尽管新兴的波斯语NLP基准已拓展至语用和礼貌范畴,却很少区分记忆的文化事实与对隐含社会规范的推理能力。我们提出DivanBench,一个聚焦迷信与习俗的诊断性基准,涵盖315道题,涉及三类任务:事实检索、配对情景验证和情境推理。评估7个波斯语大模型后发现三大缺陷:多数模型存在严重盲从偏差,能识别正确行为但无法拒绝明显违规;持续波斯语预训练非但未提升推理能力,反而加剧偏差,削弱对矛盾的辨识力;所有模型在事实检索与情境应用间存在21%的性能差距。结果表明,文化理解不能仅靠单语数据规模扩展,当前模型仅模仿文化模式,未内化其底层结构。

原文摘要 · Abstract (English)

While emerging Persian NLP benchmarks have expanded into pragmatics and politeness, they rarely distinguish between memorized cultural facts and the ability to reason about implicit social norms. We introduce DivanBench, a diagnostic benchmark focused on superstitions and customs, arbitrary, context-dependent rules that resist simple logical deduction. Through 315 questions across three task types (factual retrieval, paired scenario verification, and situational reasoning), we evaluate seven Persian LLMs and reveal three critical failures: most models exhibit severe acquiescence bias, correctly identifying appropriate behaviors but failing to reject clear violations; continuous Persian pretraining amplifies this bias rather than improving reasoning, often degrading the model's ability to discern contradictions; and all models show a 21\% performance gap between retrieving factual knowledge and applying it in scenarios. These findings demonstrate that cultural competence requires more than scaling monolingual data, as current models learn to mimic cultural patterns without internalizing the underlying schemas.

语言模型文化推理波斯语评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。