用提示工程提升低资源方言翻译,解决大模型在斯利赫蒂语中的词汇偏差问题。
LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
- 构建含2260词核心词表的上下文提示框架,嵌入语言规则与真实性校验
- 在双语方向上使翻译质量显著提升,人工评估显示幻觉减少57%
- 适合低资源方言翻译研究者,可直接复现于其他濒危语言场景
大型语言模型(LLMs)通过提示已展现出强大的翻译能力,但其在方言及低资源环境下的表现仍不明确。本研究首次系统性地考察了基于LLM的斯利赫蒂语(一种低资源孟加拉语方言)机器翻译。评估了五种先进LLM(GPT-4.1、LLaMA 4、Grok 3、DeepSeek V3.2),发现其在方言词汇处理上表现不佳。为此提出Sylheti-CAP(上下文感知提示)框架,包含三步:嵌入语言规则书、2,260个核心词汇与习语词典、真实性检查机制。大量实验表明,该框架在不同模型和提示策略下均能持续提升翻译质量。自动指标与人工评估均证实其有效性,定性分析显示幻觉、歧义和生硬表达显著减少,确立了Sylheti-CAP在方言与低资源机器翻译中的可扩展性。数据集链接:https://github.com/TabiaTanzin/LLMs-for-Low-Resource-Dialect-Translation-Using-Context-Aware-Prompting-A-Case-Study-on-Sylheti.git
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated strong translation abilities through prompting, even without task-specific training. However, their effectiveness in dialectal and low-resource contexts remains underexplored. This study presents the first systematic investigation of LLM-based machine translation (MT) for Sylheti, a dialect of Bangla that is itself low-resource. We evaluate five advanced LLMs (GPT-4.1, GPT-4.1, LLaMA 4, Grok 3, and DeepSeek V3.2) across both translation directions (Bangla $\Leftrightarrow$ Sylheti), and find that these models struggle with dialect-specific vocabulary. To address this, we introduce Sylheti-CAP (Context-Aware Prompting), a three-step framework that embeds a linguistic rulebook, a dictionary (2{,}260 core vocabulary items and idioms), and an authenticity check directly into prompts. Extensive experiments show that Sylheti-CAP consistently improves translation quality across models and prompting strategies. Both automatic metrics and human evaluations confirm its effectiveness, while qualitative analysis reveals notable reductions in hallucinations, ambiguities, and awkward phrasing, establishing Sylheti-CAP as a scalable solution for dialectal and low-resource MT. Dataset link: \href{https://github.com/TabiaTanzin/LLMs-for-Low-Resource-Dialect-Translation-Using-Context-Aware-Prompting-A-Case-Study-on-Sylheti.git}{https://github.com/TabiaTanzin/LLMs-for-Low-Resource-Dialect-Translation-Using-Context-Aware-Prompting-A-Case-Study-on-Sylheti.git}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。