用文化背景增强大模型推理,提升小语种任务表现
Culturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource Languages
- 结合文化向量检索与推理链,增强模型文化理解
- 在约鲁巴谚语解析中准确率显著优于传统方法
- 揭示翻译指标与文化相关性间的巨大差距
大语言模型在低资源语言的文化特异性推理任务中表现不佳,限制了其全球应用。为填补这一空白,本文提出文化根基的思维链(CG-CoT) prompting 方法,融合密集向量检索的文化背景信息与显式推理序列。在约鲁巴谚语解读任务上的大量实验表明,CG-CoT 在文化契合度和推理深度上均显著优于传统 prompting 方法,经自动化指标与基于 LLM 的评估双重验证。值得注意的是,我们发现词级翻译指标(如 BLEU)与人工评判的文化相关性之间存在显著差异,提示需重新思考低资源自然语言处理的评估方式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) struggle with culturally-specific reasoning tasks, particularly in low-resource languages, hindering their global applicability. Addressing this gap is crucial for equitable AI deployment. We introduce Culturally-Grounded Chain-of-Thought (CG-CoT), a novel prompting strategy that combines dense vector retrieval of cultural context with explicit reasoning sequences. Our extensive experiments on Yoruba proverb interpretation demonstrate that CG-CoT provides significantly higher culturally-aligned accuracy and depth than traditional prompting methods, validated through both automated metrics and LLM-based evaluations. Notably, we uncover stark disparities between token-level translation metrics like BLEU and human-judged cultural relevance, suggesting a rethinking of evaluation approaches for low-resource NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。