构建首个针对孟加拉语的指令微调数据集,解决低资源语言生成中的礼貌用语与结构错配问题。
Polite on the Surface, Broken in Practice: A Curated Dataset for Fixing Generation and Register Failures in Low-Resource Bangla Text Generation

- 构建4196对精心标注的孟加拉语对话交互对,聚焦文化语用与敬语一致性。
- 在4位浮点量化下使用LoRA微调模型,显著提升生成内容的语法准确性和敬语匹配度。
- 适合低资源语言研究者、多语种对话系统开发者及文化敏感性文本生成方向人士。
多语言大模型虽显著提升了跨语言对话能力,但在处理低资源语言如孟加拉语时,仍面临文化语用和上下文依赖沟通的严重瓶颈。现有先进模型在结构变化、地域习语和敬语一致性方面表现不佳。为此,我们提出一个名为「BangLa Application and DialoguE generation - BLADE」的新颖、文化对齐的指令微调数据集及评测框架,包含4196个精心标注的交互对。利用该资源,我们系统地微调并评估主流开源模型(如DeepSeek-8B和LLaMA-3.2-3B),采用参数高效微调方法LoRA,并在4位浮点数(NF4)量化框架下进行训练。实证结果表明,基于本数据集微调的模型在结构保真度和敬语一致性方面均有显著提升,为弥合低资源多语言文本生成中的语用差距提供了严谨基准。代码与数据集:https://github.com/ashuvo25/Bangla_Application_LLM/tree/main
原文摘要 · Abstract (English)
Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling culturally nuanced and context-dependent communication remains a critical bottleneck. Specifically, existing state-of-the-art models exhibit a severe pragmatic gap when handling structural variations, regional idioms, and honorific consistencies in low-resource contexts like Bangla. To address this limitation, we introduce a novel, culturally aligned instruction-tuning dataset for \textbf{BangLa Application and DialoguE generation - BLADE} and benchmarking framework comprising $4,196$ meticulously curated interaction pairs. We leverage this resource to systematically fine-tune and evaluate leading open-weight architectures, including DeepSeek-8B and LLaMA-3.2-3B, utilizing parameter-efficient fine-tuning via LoRA adapters in a 4-bit NormalFloat (NF4) quantization framework. Our empirical evaluations demonstrate that models fine-tuned on our dataset yield substantial improvements in structural fidelity and honorific alignment, providing a rigorous benchmark for bridging pragmatic disparities in low-resource multilingual text generation. Code and dataset: https://github.com/ashuvo25/Bangla_Application_LLM/tree/main
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。