构建首个孟加拉语反事实问答数据集,解析模型依赖记忆还是上下文。
Understanding QA generation: Extracting Parametric and Contextual Knowledge with CQA for Low Resource Bangla Language
- 构建反事实问答数据集BanglaCQA,支持分析模型知识来源。
- 链式思维提示使解码器型大模型在反事实场景中更有效提取记忆知识。
- 适用于低资源语言的模型可解释性研究,尤其关注孟加拉语场景。
针对孟加拉语等低资源语言的问答(QA)模型,受限于标注数据少和语言复杂性,难以判断其生成答案时更依赖预编码的参数化知识还是上下文输入。现有孟加拉语QA数据集缺乏结构化设计以支持此类分析。本文提出首个孟加拉语反事实问答数据集BanglaCQA,通过扩展原数据集并引入反事实段落与可回答性标注实现。同时,构建针对编码器-解码器语言模型、多语言基线模型的微调管道,以及面向解码器仅模型的提示方法,用于在事实与反事实场景下分离参数化与上下文知识。采用基于大模型与人工评估的方法衡量答案语义相似度。详细分析不同问答设置下的模型表现,发现链式思维(CoT)提示在反事实场景中显著提升解码器仅模型提取参数化知识的能力。本工作不仅建立分析孟加拉语问答知识来源的新框架,还揭示了低资源语言中反事实推理的关键方向。
原文摘要 · Abstract (English)
Question-Answering (QA) models for low-resource languages like Bangla face challenges due to limited annotated data and linguistic complexity. A key issue is determining whether models rely more on pre-encoded (parametric) knowledge or contextual input during answer generation, as existing Bangla QA datasets lack the structure required for such analysis. We introduce BanglaCQA, the first Counterfactual QA dataset in Bangla, by extending a Bangla dataset while integrating counterfactual passages and answerability annotations. In addition, we propose fine-tuned pipelines for encoder-decoder language-specific and multilingual baseline models, and prompting-based pipelines for decoder-only LLMs to disentangle parametric and contextual knowledge in both factual and counterfactual scenarios. Furthermore, we apply LLM-based and human evaluation techniques that measure answer quality based on semantic similarity. We also present a detailed analysis of how models perform across different QA settings in low-resource languages, and show that Chain-of-Thought (CoT) prompting reveals a uniquely effective mechanism for extracting parametric knowledge in counterfactual scenarios, particularly in decoder-only LLMs. Our work not only introduces a novel framework for analyzing knowledge sources in Bangla QA but also uncovers critical findings that open up broader directions for counterfactual reasoning in low-resource language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。