arXiv:2606.01879cs.CL2026-06

测试大模型在真实文化情境中的推理能力,发现知识多不代表用得好。

CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs

论文配图:CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
图 1 · 摘自论文原文
  • 基于原子文化规范设计可验证的评测任务
  • 顶级模型在开放问答中表现大幅下降,跨区域差异明显
  • 适合关注AI文化偏见与公平性的研究者

现有研究将大模型的文化智能简化为知识问题,忽略了其在真实场景中运用知识的能力。为此,我们提出CultureForest,一个面向文化规范根基推理的基准评测。每个问题基于一组原子规范,支持从选择题到开放式生成的渐进式评估。该基准涵盖53个地区、8个领域,共5,378个样本。大量实验表明,即使顶尖模型在开放生成任务中表现显著下滑,且存在明显跨区域差异。分析揭示:(1)推理阶段提升有限,可能加剧不平等;(2)模型表现出高度一致的区域偏好结构;(3)在严格文化约束下反应极为保守;(4)分离文化知识获取与推理后发现,尽管模型具备丰富文化知识,但有效使用仍是瓶颈。这些结果提示应从知识中心转向知识驱动的推理评估。

原文摘要 · Abstract (English)

Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their acquired knowledge in realistic scenarios. To bridge this gap, we introduce CultureForest, a benchmark for \textit{Cultural Norm Grounded Reasoning}. Each question is grounded in a small set of atomic norms, enabling verifiable and attributable evaluation. CultureForest comprises 5,378 examples across 8 domains and 53 countries/regions, and supports a progressive evaluation from multiple-choice to open-ended generation. Extensive experiments reveal that even top-tier models degrade substantially in open-ended settings, accompanied by pronounced cross-region disparities. Through targeted analysis, we uncover several consistent patterns: (1) test-time reasoning yields limited gains and may exacerbate inequity; (2) models exhibit highly shared regional preference structures; (3) model responses are markedly conservative, especially under stricter cultural constraints; and (4) by disentangling cultural knowledge acquisition from cultural reasoning, we show that while LLMs possess substantial cultural knowledge, their performance is further bottlenecked by its effective use. These findings point to a necessary shift from knowledge-centric evaluation toward measuring knowledge-grounded reasoning.

文化推理大模型评估偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。