arXiv:2602.18429cs.CLcs.IR2026-02被引 2

构建印度文化多跳问答数据集,提升大模型跨文化推理能力。

VIRAASAT: Traversing Novel Paths for Indian Cultural Reasoning

  • 基于700+文化实体知识图谱,自动生成跨州多跳问题。
  • 3200+题目验证大模型在文化推理中难以融合低概率事实。
  • 提出符号化推理框架SCoM,性能比传统方法高20%。

大语言模型在数学、编程等任务上表现优异,但在需要丰富社会文化知识和本地语境的任务中表现下降,尤其在涉及印度文化时。现有文化评测数据集存在三方面缺陷:人工构建、单跳问题为主、难以扩展。为此,我们提出VIRAASAT,一种半自动化多跳问答数据集生成方法,覆盖印度28个邦和8个中央直辖区,基于包含700多个专家标注文化实体的知识图谱,涵盖13类文化属性(如历史、节日等),生成超过3200个需链式推理的多跳问题。我们在VIRAASAT上评估当前SOTA LLMs,发现其在细粒度事实整合与因果推理上存在明显短板,即使在思维链(CoT)微调后仍无法有效建模低概率事实。为此,我们提出符号化链式操作(SCoM)框架,通过模拟知识图谱内部原子操作训练模型,使其能可靠地遍历图结构拓扑。监督微调实验表明,SCoM相较标准CoT基线最高提升20%。我们公开VIRAASAT数据集与成果,为构建文化感知推理模型奠定基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant progress in reasoning tasks across various domains such as mathematics and coding. However, their performance deteriorates in tasks requiring rich socio-cultural knowledge and diverse local contexts, particularly those involving Indian Culture. Existing Cultural benchmarks are (i) Manually crafted, (ii) contain single-hop questions testing factual recall, and (iii) prohibitively costly to scale, leaving this deficiency largely unmeasured. To address this, we introduce VIRAASAT, a novel, semi-automated multi-hop approach for generating cultural specific multi-hop Question-Answering dataset for Indian culture. VIRAASAT leverages a Knowledge Graph comprising more than 700 expert-curated cultural artifacts, covering 13 key attributes of Indian culture (history, festivals, etc). VIRAASAT spans all 28 states and 8 Union Territories, yielding more than 3,200 multi-hop questions that necessitate chained cultural reasoning. We evaluate current State-of-the-Art (SOTA) LLMs on VIRAASAT and identify key limitations in reasoning wherein fine-tuning on Chain-of-Thought(CoT) traces fails to ground and synthesize low-probability facts. To bridge this gap, we propose a novel framework named Symbolic Chain-of-Manipulation (SCoM). Adapting the Chain-of-Manipulation paradigm, we train the model to simulate atomic Knowledge Graph manipulations internally. SCoM teaches the model to reliably traverse the topological structure of the graph. Experiments on Supervised Fine-Tuning (SFT) demonstrate that SCoM outperforms standard CoT baselines by up to 20%. We release the VIRAASAT dataset along with our findings, laying a strong foundation towards building Culturally Aware Reasoning Models.

文化推理知识图谱多跳问答LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。