针对低资源语言的文化常识问答,提出区域感知混合检索方法。
Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

- 结合词法匹配与语义相似度,加入区域权重提升检索相关性。
- 在30种语言的BLEnD数据集上,跨语言稳定性显著优于纯参数推理。
- 适合多语言文化常识问答研究者,尤其关注低资源语言场景。
尽管大型语言模型在通用领域推理任务中表现优异,但在数字和文本数据较少的语言中,面对基于文化的知识时仍面临挑战。本文以包含30种语言、涵盖饮食、体育、家庭等社会文化领域的BLEnD基准为测试平台,研究文化常识多选题问答。提出一种区域感知的混合检索方法,融合BM25词法匹配与密集语义相似度,并引入区域加权启发式策略以提高答案相关性。检索到的文档用于构建结构化提示,输入至量化版Qwen3-14B模型,采用基于逻辑值的确定性答案选择。实验表明,该混合检索方法在跨语言稳定性上优于纯参数推理;但高资源与低资源语言间仍存在明显性能差距,说明检索增强方法无法完全克服训练数据不平衡带来的限制。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, they may face challenges with culturally grounded knowledge within languages with limited digital and textual data. In this paper, we investigate culturally grounded multiple-choice question answering with the BLEnD benchmark, which consists of a multilingual corpus of 30 languages and covers various socio-cultural domains, such as cuisine, sports, family, etc. We propose a region-aware hybrid retrieval approach that combines BM25 lexical matching and dense semantic similarity with regional weighting heuristics to improve the relevance of the answer. The retrieved documents are used to construct a structured prompt for the Qwen3-14B quantized model with logit-based deterministic answer selection. The experimental results show improvements to cross-lingual stability with the hybrid retrieval approach over pure parametric inference for culturally grounded question answering. However, there are still notable performance gaps between languages with more and less training data. This shows that the limitations of the retrieval augmentation approach are not entirely overcome by the training data imbalance problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。