构建首个印尼文化多跳问答数据集,测试大模型真实文化推理能力
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
- 将单跳文化问题转化为六类线索的多跳推理链
- 评估显示主流模型在复杂文化推理上仍有显著差距
- 适合研究大模型文化理解与跨文化AI的学者使用
理解文化需要跨越语境、传统和隐含社会知识的推理,远超于孤立事实的回忆。然而,大多数文化相关的问答基准依赖单跳问题,可能使模型利用浅层线索而非展现真正的文化推理能力。本文提出ID-MoCQA,首个基于印尼传统的大规模多跳问答数据集,支持英、印尼双语。我们设计新框架,系统地将单跳文化问题转化为涵盖六类线索(如常识、时间、地理)的多跳推理链。通过专家评审与大模型作为裁判的多阶段验证流程,确保问答对质量。在先进模型上的评估显示,文化推理能力存在显著短板,尤其在需要细微推断的任务中。ID-MoCQA为提升大模型的文化素养提供了挑战性且关键的基准。
原文摘要 · Abstract (English)
Understanding culture requires reasoning across context, tradition, and implicit social knowledge, far beyond recalling isolated facts. Yet most culturally focused question answering (QA) benchmarks rely on single-hop questions, which may allow models to exploit shallow cues rather than demonstrate genuine cultural reasoning. In this work, we introduce ID-MoCQA, the first large-scale multi-hop QA dataset for assessing the cultural understanding of large language models (LLMs), grounded in Indonesian traditions and available in both English and Indonesian. We present a new framework that systematically transforms single-hop cultural questions into multi-hop reasoning chains spanning six clue types (e.g., commonsense, temporal, geographical). Our multi-stage validation pipeline, combining expert review and LLM-as-a-judge filtering, ensures high-quality question-answer pairs. Our evaluation across state-of-the-art models reveals substantial gaps in cultural reasoning, particularly in tasks requiring nuanced inference. ID-MoCQA provides a challenging and essential benchmark for advancing the cultural competency of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。