arXiv:2607.17828cs.CL2026-07

解决孟加拉语同形异义词的文化歧义问题,提升低资源模型理解力。

When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs

论文配图:When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs
图 1 · 摘自论文原文
  • 构建1516句专家标注的孟加拉语同形异义句数据集,含3032个标注实例
  • 发现主流模型普遍倾向常见词义,文化名称识别错误率高达100%
  • 通过对比链式思维提示和文化解释蒸馏,小模型准确率显著提升

许多孟加拉语词汇既是个人姓名又是富含文化内涵的普通名词,如“Maya”既为女性名,也指深情的怜爱。正确解读需依赖训练数据中稀缺的文化知识。本文提出文化纠缠同形异义(CEH)消歧任务,构建包含1,516句、3,032个标注实例的孟加拉语基准数据集,每例均标注文化语境类别及推理依据。实验表明,各类开源与闭源模型普遍存在主导词义偏差:倾向于选择常见名词义而忽略人名。特定孟加拉语模型在所有提示策略下表现失败,说明仅靠语言专项预训练不足以获得文化理解。进一步发现,对比链式思维提示可无须训练显著降低偏差;将文化解释进行蒸馏后,1-3B参数的小模型能学会推理而非死记硬背标签,使主导词义偏差从最高100%降至不足5%,并使原本失败的孟加拉语模型成为最强系统。数据集与代码已公开于https://github.com/ashuvo25/BanglaCEH。

原文摘要 · Abstract (English)

Many Bangla words are at once personal names and culturally loaded common nouns, "Maya" is both a girl's name and a word for affectionate compassion. Choosing the right reading demands cultural knowledge that is scarce in the pretraining data of modern language models. We introduce Culturally Entangled Homograph (CEH) disambiguation and build a Bangla benchmark of 1,516 expert-verified sentences (3,032 labelled occurrences) in which one word appears twice with two distinct readings, each labelled with a culturally grounded category and an explanation of the reasoning behind it. Across open- and closed-source models, we find a systematic dominant-meaning bias: models default to the common-noun sense and overlook the name. A Bangla-specific model fails under every prompting regime we test, showing that language-specific pretraining alone does not confer cultural grounding. We further show that contrastive chain-of-thought prompting can sharply reduce this bias without training, and that distilling cultural explanations teaches small (1-3B) models to reason toward the correct reading rather than memorise labels, cutting dominant-meaning bias from as high as 100% to under 5% and turning the failed Bangla-specific model into our strongest system. Dataset and code are available at https://github.com/ashuvo25/BanglaCEH.

文化理解同形异义低资源模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。